In a New Test, AI Mortgage Assistants Got Nearly 1 in 4 Answers Wrong
When researchers asked three of the leading general-use AI models which deposits in a bank statement “could be of foreign origin,” transactions connected to English names were flagged 13.3% of the time. But when the deposit came from a non-English name, that number jumped to 77%.
“That's not how America works,” Matthew Toles, a Columbia University doctoral student and co-author of the new study, “MortarBench: Evaluating Mortgage Loan Origination Agents,” tells Realtor.com®. “You aren't determined whether you're foreign-based on how your name sounds.”
Even as concerns persist about AI’s ability to provide accurate and unbiased information, mortgage lenders are moving to adopt it. More than 80% were evaluating the technology as of June, according to a survey by The Mortgage Collaborative, while 17% had deployed it in live production workflows.
Now, Toles and his colleagues' work is giving the industry a better way to test it.
“Everybody [is] using AI, but nobody really understands how to use AI in compliance and with the correct guardrail,” says Diane Yu, co-founder and CEO of Tidalwave, a mortgage technology company that collaborated with Columbia researchers on the work.
In financial services, she says, “the key difference is not just about utilizing AI,” but using it correctly.
Where top AI models struggled the most
In their new study, the researchers introduce MortarBench, an open-source benchmark that lets companies measure AI against the same set of mortgage origination tasks. The idea is to give lenders, regulators, and even borrowers a common understanding of how accurate a model is.
To build the benchmark, researchers drew from real questions submitted to a mortgage assistant, then narrowed them to the most common and useful in loan origination: Do payroll deposits match the employer listed on the application? Which deposits are large enough to require scrutiny? Is an account jointly held with someone who isn't applying for the mortgage?
That kind of repetitive, detail-heavy work may seem well suited to AI, but even the strongest general-purpose models tested didn't always get the entire answers right.
On the benchmark's strictest measure—whether the complete answer matched the known correct one—Gemini 3.1 Pro was correct 77.1% of the time, GPT-5.5 76.8%, and Claude Sonnet 4.6 51.4%.
A particularly revealing weakness was picking specific transactions out of a bank statement.
Zhou Yu, an associate professor at Columbia and study co-author, compared the job to finding “the needle in the haystack.”
“It's like a needle. You find the needle in the haystack,” she says. “You have so many transactions; it's very easy to miss one or two.”
But models often made the opposite mistake, too, pulling in transactions that didn't belong.
When researchers manually reviewed Gemini's incorrect answers on transaction-list questions, the most common problem was misclassification.
In one case, Gemini counted a personal loan as a buy-now-pay-later transaction. Other errors included assuming all wire transfers were international, treating deposits from co-borrowers as automatically documented, and classifying a one-time housing payment as recurring.
So the problem wasn't simply finding the needle—it was reliably knowing what counted as one.
But Toles cautions these scores were for what he called “naive use of foundational models, largely equivalent to taking the application package, pasting it into ChatGPT, and asking it a bunch of questions about it.”
Major industry players, he adds, typically develop their own proprietary models that perform better on these tasks. Tidalwave, for example, scored 95% on yes or no questions in a separate and earlier benchmark test.
Why mortgage lenders need a common AI test
But understanding that gap—between naive use and specialized models—is exactly the standardization that the industry may need, as mortgage lenders are under intense pressure to make an expensive, labor-heavy process faster.

Originating a retail mortgage cost lenders about $11,800 per loan in the second quarter of 2025, according to Freddie Mac. Meanwhile, adopting its basic digital underwriting capabilities averaged about $1,700 in savings per loan and production times that were five days shorter.
Through that lens, it's easy to understand the industry's appetite for AI and the incentive to adopt any model available. But getting the work done faster is only useful if it is also done correctly.
Mortgage lending is one of the most heavily regulated corners of consumer finance, and Fannie Mae and Freddie Mac formalized that concern in 2026 with new AI governance requirements for their seller-servicers.
The rules put the impetus on companies to manage risks from AI, including overseeing systems supplied by outside vendors. They must also, when requested, disclose what AI they use, how they use it, and what safeguards are in place.
For an industry looking to cut down on burdensome work, it's a lot of new and onerous responsibilities to take on. And as lenders race toward automation, they must also answer the thorny question: How do they know those safeguards actually work?
What the new benchmark can tell borrowers
The answers could matter outside the industry, too. As users feed more sensitive financial and personal identifying information to AI models, errors and hidden biases can carry higher stakes.
In one 2025 study of U.S. ChatGPT users, more than a third said they had discussed their personal finances with the chatbot—even as 82% described their AI conversations as sensitive or highly sensitive
“We don't have visibility into what these models are doing, what companies are doing with them, and what the outcomes are,” Toles warns.
A mortgage application is just one example. It's rife with details about income, debts, account balances, and individual bank transactions.
That's why Yu says borrowers should ask lenders whether that information is being passed to an outside large language model and whether their AI has undergone independent evaluation.
“You should ask those questions,” she says. “You should be very careful.”
The MortarBench study helps quantify those concerns. And because it's open source, those same questions can now be put to other AI systems rather than leaving each lender or vendor to define success for itself.
Toles says that will become especially important if many mortgage companies rely on the same underlying models.
“Supposing everybody is using the same models or using them in similar ways, and we see like we have established that there are systemic biases in how these models behave,” he says. “Is this possibly going to create some systemic risk across the industry?”
In his words, “If we don't measure it, then we don't know about it."
Recent Posts










"My job is to find and attract mastery-based agents to the office, protect the culture, and make sure everyone is happy! "
