Skip to main content
For decades, market intelligence sold trust through reputation. You trusted the report because you trusted the name on the cover. That kind of trust doesn’t survive the handoff to a machine. An agent doesn’t buy a name; it only trusts what it can verify. So that’s what we built. Gildea follows the AI race as it happens and hands your agents structured intelligence with receipts attached. Every sentence we serve is grounded in its source and verified before it reaches you, giving your agent trustworthy intelligence it can build on. The proof is in the numbers we hold ourselves to:
>99%
of what we serve carries no material distortion
0
zero tolerance for a wrong number, date, or fact
7
kinds of error checked, not just false facts
“No material distortion” is stricter than “no hallucinated facts.” It also rules out subtle hallucinations, the quieter distortions that leave a fact intact but still change what a reader concludes. We run a standing evaluation of what we serve and publish the method so you can run it yourself. Here is what happens before a unit reaches you, and how we hold ourselves to that number.

Grounding, then verification

Every unit Gildea serves pairs one claim or sentence with the evidence that backs it. Before that pair is served, two things must be true:
  1. Grounded. The evidence appears word-for-word in the source.
    • Evidence we can’t match verbatim is treated as a hallucinated quote and discarded; the unit survives only if real, verbatim evidence backs it.
  2. Verified. That pairing runs through a loop built to break the match between the text and the underlying evidence.
1. Verify. Each unit is scored against its evidence: claims to a strict entailment standard, sentences to a faithfulness standard. Then a battery of deterministic checks runs on every pair, regardless of score: contradictions, negation flips, quantity and date mismatches, misquotes, missing entities, and a bidirectional check that catches a claim or sentence quietly overstating what its evidence supports. These checks can only downgrade a pass, never rescue a fail. When in doubt, the unit doesn’t pass. 2. Repair. A weak unit isn’t dropped on the first try. The system hunts the source for better verbatim evidence and keeps it only if the score actually improves. If better evidence isn’t enough, it rewrites the unit and runs it back through the entire loop. Repair before rejection. 3. Triage. What survives in the gray zone gets adjudicated by independent model judges, blind. A unit is only cleared when they fully agree. 4. Human review. Only the genuine judgment calls reach a person, never the easy passes. Reviewers can override a verdict, correct the evidence and trigger a re-score, or fix a theme. Every override carries an audit trail. The result is the contract that matters: only verdict=pass data is served, and presence is the verdict. If a unit is in the response, it earned its way there.

Trust contract

What you get back is the verified part of a source, not all of it. The claims and sentences in a signal are the ones that cleared the loop; the rest were pruned. Read a signal as Gildea’s verified view of a source, not a transcript of it. Thesis and synopsis text is the exception: it’s always served whole.

What the API exposes

By default a unit carries no verification block at all. It doesn’t need one: being served already means it passed. The block shows up in exactly one case, when a human reviewer signed off.
That’s the whole contract. Raw scores, scoring modes, thresholds, and reason codes stay internal. They’re the machinery used to decide pass or fail, not something you need to read. What you get instead is the part that lets you check the work yourself: every unit ships with its evidence_preview, and every search result ships with a citation.

How we evaluate our output

The loop decides pass or fail on every unit. To measure how well it does that, we run a standing evaluation of what we serve. Most faithfulness evaluations check one thing: did the model state a false fact. Ours checks seven kinds of error. Six of them are subtle hallucinations, the ones that survive a citation check because the quote is real and the sentence still misleads. Every check compares one served sentence against its own source article. Nothing is judged across articles or against outside knowledge. The evaluation reads a sample of served units. We cap how many come from any one article, spread the sample across unit types, and weight it to match the production mix, so no single article or type can skew the result. It runs on a rolling schedule:
  1. Read against the source. Judge each sampled unit against its full source, on all seven error kinds.
  2. Independent judge. The judge is a different model family from the one that runs verification, so it can’t share its blind spots.
  3. Calibrate the judge. Before any real verdict is read, the judge has to pass planted items: obvious cases, hard cases already settled by a person, and the same case twice for consistency. We also plant real errors disguised as harmless ones, to measure what the judge misses.
  4. Materiality. An error only counts if it would change what a reader concludes. Disputed calls get a second independent judge, and a person makes the final call.
  5. Pre-register. The pass thresholds, and the statistics that test them, are fixed before each run, so no one grades to the result.
The headline number is a 95% confidence bound, not a point estimate. A wrong number, date, or fact is a separate, zero-tolerance bar. One kind of error this evaluation can’t catch, we catch elsewhere: a unit whose text is faithful but whose entity tag points at the wrong company or person. That is handled by entity resolution, not by reading the source.

Don’t trust us. Verify us.

The evidence_preview we ship is a short verbatim excerpt, capped at 50 characters. It is deliberately a preview, not the evidence itself: it shows where the unit is grounded, and citation.url is where you verify. We cap it on purpose: Gildea points to a source, it does not reproduce or redistribute it. That is a fair-use line we hold. Use the snippet as a pointer. Paired with its citation, it leads back to the full source, and that is where verification actually happens, ours and yours. Our evaluation reads served units against those same public sources. Nothing it depends on is anything we keep to ourselves. So don’t take our verdict on faith. Run the evaluation yourself, on any slice of Gildea, against the sources we cite. Find one we missed. We want it.

Recipe: Run the evaluation yourself

Follow a unit’s citation back to the source and check our verdict, on any slice of Gildea.