Silvia Beats Every Major AI Product on Credit Cards

Learn how Silvia outperformed major AI products on difficult credit card questions.
We took five of the most popular AI products and put them through ten of the hardest credit card questions we could write. The kind with year-specific figures, terms that interact in strange ways, and issuer-specific rules that break from the defaults. Every answer was graded from 1 to 10 for factual accuracy by a judge who did not know which product produced it.
Silvia scored 8.60, the highest of any product tested. The next-best came in at 7.40, and the rest of the field trailed from there.
That result is the headline. The more interesting story is why it happened, and we ran a much larger study to answer that, which we will get to below.
Why Silvia comes out ahead
Most AI tools answer credit card questions from memory. They absorbed a lot of text during training, and when you ask a question, they produce a fluent answer based on what they remember. That works fine for common questions and falls apart on the hard ones, where the answer turns on the exact wording of an issuer's terms or a figure that changed this year.
Silvia works differently. We built a curated library of primary consumer-credit rules: the Truth in Lending Act and Regulation Z, the CFPB's rulemakings and guidance, the operating rules of the card networks, and the terms and disclosures issuers publish for their cards. When you ask Silvia a credit card question, it retrieves from that library at the moment you ask, and it can run a live web search when it needs current information. Instead of hoping the AI remembers the rules correctly, we give it the actual rules to work from, and we make it cite the source so you can check the answer yourself.
That is the difference the chart is measuring. On hard credit card questions, retrieving from the real rules beats answering from memory.
We went deeper than ten questions
Ten questions make a clean headline, but they cannot tell you how much the tooling matters or when. So we ran a second, much larger study: 100 realistic advisory scenarios, each pairing a user's financial profile with a question about their card portfolio.
We tested the same system two ways: as the base model with no tools, and with the full Silvia harness. That let us isolate exactly what the tooling adds. We also scored every answer on two separate things. Accuracy: is the answer correct? Grounding: can you verify it, meaning does it cite the real card terms, and when you look that citation up, does it actually say what the answer claims? An answer can be perfectly correct and still useless if there is no way to check it, which is why we scored the two separately.
What the deeper study found
The tooling made a large difference. The full harness beat the base model by a wide margin on accuracy and a wider one on grounding.
The pattern was consistent. The base model often produced a plausible answer that could not be checked, naming cards and quoting benefits from memory with no way to verify any of it. When Silvia retrieved from the actual card terms, the answer either matched the source or the source made clear where the answer needed to change.
That distinction matters more than any single number. An answer a cardholder cannot check against the real terms is not one they can act on, even when it happens to be right. The library is what turns a plausible recommendation into one worth acting on.
We published all of it
We put all 100 questions and their scores online under an open license, free for anyone to inspect. If we are going to say Silvia is the most accurate AI on credit cards, the right thing to do is let people check that claim against the actual data rather than take our word for it.
There is a broader reason too. Generic AI benchmarks tell you which model is smartest in general. They do not tell you whether an AI system actually works in a specific field like credit cards, where the answer has to match the real terms, for the real year, with a citation you can follow. The only way to know that is to test it in the field itself and show your work.
Explore the full benchmark here: huggingface.co/datasets/cfosilvia/silvia-creditcard-bench
Financial information notice
This content is for informational purposes only. It is not financial, investment, or legal advice. Past performance does not guarantee future results. Consult a qualified professional before making financial decisions.
Share this article
Your money deserves superintelligence.
Free forever. No card required. Give Silvia five minutes and see what an AI CFO trained on your money actually knows.


