Bring one document set and one workflow. We will run this against your material and show you the same page with your numbers on it.
Talk to Us →The value is in what survives the trip. An answer that loses the table it depended on, misattributes a clause, or cannot say which page it came from is worse than no answer — it is an answer someone will act on. So we publish what survived, what did not, and how we counted.
Measured on the ServAI benchmark corpus · FMSC-Model-Ed4.pdf · PMSA-Service-Agreement.pdf · MS-114-Routine-Inspection.pdf · Q2-Q3-Maintenance-Register.xlsx · 140 pages · seed 20260803
Sections were sampled at random, the source reopened, and each one re-found on the page the index recorded. A section that cannot be re-found is a failure and is named, not dropped. Same seed, same result — a customer can rerun it.
Cross-references resolved: 1,242 of 1,389 (89.4%) — a reference pointing at a clause the index does not hold is counted as unresolved rather than quietly dropped. Ingested in 28.8s.
Reconcile Q2–Q3 maintenance billing against the service agreement, and confirm inspection intervals meet the code. Run against the same four files: the model code (FMSC-Model-Ed4.pdf), the service agreement (PMSA-Service-Agreement.pdf), method statement MS-114 and a 245-row maintenance register.
Of AED 284,456 billed across the period, AED 12,118 is not payable on the terms of the agreement.
| Category | Lines | Value (AED) | Grounding | Status |
|---|---|---|---|---|
| Billed above contracted rate | 19 | 3,436 | Contract Table 7.1 · p.4 vs register | Recoverable |
| Visit invoiced twice | 5 | 6,980 | Clause 11.1 · register duplicates | Not payable |
| Completed visit without evidence (15% deduction) | 9 | 1,702 | Clause 11.1 read with 12.1 | Deductible |
| Inspection interval exceeded the code maximum | 13 | — | Code Table 3.1 · p.30 | Non-compliance |
The method statement in use permits a longer inspection interval than the code allows for the same asset classes. Both are cited. The agent did not pick one.
| Asset | Code | Source | Method statement | Source |
|---|---|---|---|---|
| AC-01 | 60 days | Code · Table 3.1 · p.30 | 90 days | MS-114 · Table 4.1 · p.2 |
| AC-03 | 60 days | Code · Table 3.1 · p.30 | 90 days | MS-114 · Table 4.1 · p.2 |
The corpus carries a known set of planted anomalies written to an answer key the investigation never opens. Comparing afterwards is what makes this falsifiable rather than illustrative.
| Category | Planted | Found | Missed | False + | Recall |
|---|---|---|---|---|---|
| Overcharged lines | 18 | 18 | 0 | 1 | 100.0% |
| Duplicate visits | 5 | 5 | 0 | 0 | 100.0% |
| Interval breaches | 13 | 13 | 0 | 0 | 100.0% |
| Missing evidence | 8 | 8 | 0 | 1 | 100.0% |
ServAI benchmark corpus · seed 20260803 · reproducible on request.
The four files are generated reference documents built to a published answer key, not customer
material and not a real code or agreement. Every figure on this page is computed from those
files and is not a measurement of a customer deployment. We will run the same audit against
your own documents under NDA.
Bring one document set and one workflow. We will run this against your material and show you the same page with your numbers on it.
Talk to Us →