>99%
of what we serve carries no material distortion
0
zero tolerance for a wrong number, date, or fact
7
kinds of error checked, not just false facts
Grounding, then verification
Every unit Gildea serves pairs one claim or sentence with the evidence that backs it. Before that pair is served, two things must be true:- Grounded. The evidence appears word-for-word in the source.
- Evidence we can’t match verbatim is treated as a hallucinated quote and discarded; the unit survives only if real, verbatim evidence backs it.
- Verified. That pairing runs through a loop built to break the match between the text and the underlying evidence.
verdict=pass data is served, and presence is the verdict. If a unit is in the response, it earned its way there.
Trust contract
What you get back is the verified part of a source, not all of it. The claims and sentences in a signal are the ones that cleared the loop; the rest were pruned. Read a signal as Gildea’s verified view of a source, not a transcript of it. Thesis and synopsis text is the exception: it’s always served whole.What the API exposes
By default a unit carries no verification block at all. It doesn’t need one: being served already means it passed. The block shows up in exactly one case, when a human reviewer signed off.evidence_preview, and every search result ships with a citation.
How we evaluate our output
The loop decides pass or fail on every unit. To measure how well it does that, we run a standing evaluation of what we serve. Most faithfulness evaluations check one thing: did the model state a false fact. Ours checks seven kinds of error. Six of them are subtle hallucinations, the ones that survive a citation check because the quote is real and the sentence still misleads.
Every check compares one served sentence against its own source article. Nothing is judged across articles or against outside knowledge.
The evaluation reads a sample of served units. We cap how many come from any one article, spread the sample across unit types, and weight it to match the production mix, so no single article or type can skew the result. It runs on a rolling schedule:
- Read against the source. Judge each sampled unit against its full source, on all seven error kinds.
- Independent judge. The judge is a different model family from the one that runs verification, so it can’t share its blind spots.
- Calibrate the judge. Before any real verdict is read, the judge has to pass planted items: obvious cases, hard cases already settled by a person, and the same case twice for consistency. We also plant real errors disguised as harmless ones, to measure what the judge misses.
- Materiality. An error only counts if it would change what a reader concludes. Disputed calls get a second independent judge, and a person makes the final call.
- Pre-register. The pass thresholds, and the statistics that test them, are fixed before each run, so no one grades to the result.
Don’t trust us. Verify us.
Theevidence_preview we ship is a short verbatim excerpt, capped at 50 characters. It is deliberately a preview, not the evidence itself: it shows where the unit is grounded, and citation.url is where you verify. We cap it on purpose: Gildea points to a source, it does not reproduce or redistribute it. That is a fair-use line we hold.
Use the snippet as a pointer. Paired with its citation, it leads back to the full source, and that is where verification actually happens, ours and yours. Our evaluation reads served units against those same public sources. Nothing it depends on is anything we keep to ourselves.
So don’t take our verdict on faith. Run the evaluation yourself, on any slice of Gildea, against the sources we cite. Find one we missed. We want it.
Recipe: Run the evaluation yourself
Follow a unit’s citation back to the source and check our verdict, on any slice of Gildea.