Silvia Labs
A New Frontier in AI Tax Intelligence: How Harness Engineering Improves Factual Accuracy and Verifiability
Does retrieval tooling make AI tax guidance more accurate, and easier to check? We put the same frontier model through 200 expert-level tax questions, with and without our retrieval harness, ran the whole study twice, and graded every answer blind against the law itself.
Silvia LabsAugust 2026200 questions1,511 answers2,949 blind judgments
200 questions
Two benchmarks of 100. The analysis rests on the 190 that both runs posed.
1,511 answers
Four tool setups, two runs. Each written in a fresh session, tool use audited from the logs.
2,949 judgments
Two rubrics per answer, graded blind. Citations looked up in the law itself.
Abstract
Language models answer tax questions fluently, but fluency is not authority. A practitioner cannot act on guidance that offers no way to check it against the law it claims to rest on. We set out to measure what retrieval tooling changes about that.
We put the same frontier model through two 100-question benchmarks under four tool configurations: once from memory alone, once with a curated library of the tax law itself, once with a web search tool it ran on its own, and once with both. The first benchmark is built to catch the model out on the fine print of the tax code. The second is made of the kind of planning questions real clients actually bring. We ran the whole study twice. Every answer was scored from 1 to 10 on two rubrics, and the judges never knew which setup had produced the answer they were grading. Four configurations, two runs, and the 190 questions both runs posed come to 1,511 answers and 2,949 blind judgments in all.
Two things we measured, and why they are separate
Accuracy
Is the answer right? The correct conclusion, the correct figures, for the correct tax year.
Grounding
Can you check it? Does the answer cite the statute or ruling it depends on, and when the judge looks that citation up, does it actually say what the answer claims?
An answer can be perfectly correct and still unusable to a practitioner if there is no way to verify it. That is why we score the two separately.
Headline findings
On adversarial statutory questions, tools carried the answer
The Silvia harness beat the model with no tools by +2.13 points on accuracy (p=0.008) and +3.40 points on grounding (p<0.001). Both runs agreed the tools helped; they disagreed on how much (+1.32 in run 1, +2.92 in run 2), so the per-run estimates carry more weight than the pooled one.
On realistic client scenarios, the model was already accurate — the library made it checkable
Accuracy barely moved with tools added (+0.13, p=0.244). What did improve was grounding (+0.38, p=0.012), and the gain came from the curated tax library rather than from web search. An answer that is correct but unverifiable is unusable to a practitioner; the library is what let them check it.
Web search and a curated library do different jobs
Web search supplies facts the model does not have. The library supplies the citation that lets someone check the answer. Across both benchmarks and both rubrics, the combined setup was the only one never significantly behind any other — the property that matters when you cannot know in advance which kind of question a user will bring.
Read the full write-up
Methodology, confidence intervals, every chart, and the limitations section — nine pages.
Data and materials
The benchmark is open. All 200 questions and their judge scores are published as silvia-tax-bench under a CC BY 4.0 license.
How to cite
Silvia Labs. A New Frontier in AI Tax Intelligence: How Harness Engineering Improves Factual Accuracy and Verifiability. Tax Intelligence Benchmark meta-analysis, August 2026. https://www.cfosilvia.com/ai-lab/tax-benchmark-analysis
Related
For a general audience
Silvia Is the Most Accurate AI on Tax. We Tested It Against Everyone.
A shorter, non-technical companion piece to this paper — what we found, why it matters, and what it means for tax professionals.
Featured on CNBC
Anthony Pompliano on Silvia AI, on Squawk Box
Silvia, Inc. chairman and CEO Anthony Pompliano joins Squawk Box to discuss Silvia's tax results and how AI is changing financial tools. Six minutes.