Back to blog

Silvia Beats Every Major AI Product on Mortgage

Bar chart from Silvia's Provider Mortgage Benchmark showing Factual Correctness Accuracy on a 1 to 10 scale. Silvia leads at 8.80, followed by Claude 8.30, Perplexity 8.10, ChatGPT 7.30, Grok 6.50, and Gemini 6.10.

Learn how Silvia outperformed major AI products on difficult mortgage questions.

We took five of the most popular AI products and put them through ten of the hardest mortgage questions we could write. The kind with year-specific figures, provisions that interact in strange ways, and state rules that break from the federal treatment. Every answer was graded from 1 to 10 for factual accuracy by a judge who did not know which product produced it.

Silvia scored 8.80, the highest of any product tested. The next-best came in at 8.30, and the rest of the field trailed from there.

That result is the headline. The more interesting story is why it happened, and we ran a much larger study to answer that, which we will get to below.

Why Silvia comes out ahead

Most AI tools answer mortgage questions from memory. They absorbed a lot of text during training, and when you ask a question, they produce a fluent answer based on what they remember. That works fine for common questions and falls apart on the hard ones, where the answer turns on the exact wording of a rule or a figure that changed this year.

Silvia works differently. We built a curated library of primary mortgage law: the Truth in Lending Act and RESPA, the FHA, VA, and USDA lender handbooks, the Fannie Mae Selling Guide, and the mortgage guidance from individual state banking regulators. When you ask Silvia a mortgage question, it retrieves from that library at the moment you ask, and it can run a live web search when it needs current information. Instead of hoping the AI remembers the rule correctly, we give it the actual rule to work from, and we make it cite the source so you can check the answer yourself.

That is the difference the chart is measuring. On hard mortgage questions, retrieving from the real rules beats answering from memory.

We went deeper than ten questions

Ten questions make a clean headline, but they cannot tell you how much the tooling matters or when. So we ran a second, much larger study: 178 expert-level mortgage questions, split into 83 everyday client conversations and 95 research-style scenarios.

We tested the same system two ways: as the base model with no tools, and with the full Silvia harness. That let us isolate exactly what the tooling adds. We also scored every answer on two separate things. Accuracy: is the answer correct? Grounding: can you verify it, meaning does it cite the real rule, and when you look that citation up, does it actually say what the answer claims? An answer can be perfectly correct and still useless to a mortgage professional if there is no way to check it, which is why we scored the two separately.

What the deeper study found

On the research-style questions, the tooling made a large difference. The full harness beat the base model by a wide margin on accuracy and a wider one on grounding.

On the everyday client conversations, something more interesting happened, and it is worth being honest about. The base model was already quite accurate on these without any tools. Adding the harness barely changed the accuracy score, because the model already handles common client questions well.

What the harness changed was verifiability. On those everyday questions, the library was the thing that let the answer cite the actual rule so a professional could check it. The accuracy was already there. The ability to trust and verify it was not, until the tooling supplied it.

That distinction matters more than any single number. An answer a loan officer cannot check against the real rule is not one they can act on, even when it happens to be right. The library is what turns a plausible answer into one worth acting on.

We published all of it

We put all 178 questions and their scores online under an open license, free for anyone to inspect. If we are going to say Silvia is the most accurate AI on mortgage, the right thing to do is let people check that claim against the actual data rather than take our word for it.

There is a broader reason too. Generic AI benchmarks tell you which model is smartest in general. They do not tell you whether an AI system actually works in a specific field like mortgage lending, where the answer has to match the real rule, for the real year, with a citation you can follow. The only way to know that is to test it in the field itself and show your work.

Explore the full benchmark here: huggingface.co/datasets/cfosilvia/silvia-mortgage-bench

Share this article

Your money deserves superintelligence.

Free forever. No card required. Give Silvia five minutes and see what an AI CFO trained on your money actually knows.