MIT study finds AI financial advice solid, but biased by gender

MIT Sloan researchers built a model of how a person's income, job, investments, and taxes typically evolve over a lifetime, then used it to test whether following AI financial advice actually helps. The team, led by assistant professor Taha Choukhmane with co-authors Weidong Lin and Matthew Akuzawa of MIT Sloan and Tim de Silva of Stanford Graduate School of Business, asked 1,000 adults to write their own prompts seeking spending and investing advice from GPT-5.2, GPT-5.6, or Gemini 3 Flash. They then simulated what would happen if people aged 22 to 89 followed that advice repeatedly over their lives, and compared the outcome both to what people already do without AI and to advice generated from detailed, academic-style prompts written by the researchers themselves.
The advice held up better than the team expected. It pushed people to save more, invest in diversified stock funds, take on less risk after age 45, and draw down savings in retirement, producing sizable saving buffers for almost everyone over 30. It fell down on subtler judgment calls: the AI tended to lean on simple rules of thumb, telling people who had just lost a job to cut spending sharply even when they had savings to draw on, and it let portfolios drift instead of actively rebalancing them. Quality improved when prompts were more structured and included information such as age, job status, income, savings, and assumptions about taxes and Social Security, but even the best prompts did not fix the rebalancing problem.
The advice also varied by who was asking. Prompts written by men, by people with higher financial literacy, or by people who had used AI before generated about 5% more wealth near retirement than prompts from other groups. Higher recommended equity allocations for men and the more financially literate compounded into roughly $50,000, or 4%, less wealth at age 60 for women and less literate users. Lower recommended saving rates for first-time AI users left them with almost $100,000, or 6%, less wealth at 60 than users with prior AI experience. About two-thirds of the gender gap traced back to men and women phrasing their questions differently (women used words like "family," "grocery," and "pay"; men used "strategy," "crypto," and "growth"), while the remaining third came from the model itself giving different advice when the same prompt was simply labeled as coming from a woman rather than a man.
The study also found the models steering people toward specific products without being asked: Vanguard investment products showed up in 6% of LLM responses and iShares in 3.4%, even though fewer than 0.4% of the original prompts named either company. Choukhmane said this could reshape how financial firms compete for attention, shifting weight away from traditional marketing and search visibility toward how favorably an LLM happens to describe a given product. The paper, "AI Financial Advice: Supply, Demand, and Life Cycle Implications," won the Swiss Finance Institute's Outstanding Paper Award 2026.
Key facts
- MIT Sloan researchers simulated lifetime finances for people aged 22 to 89 who followed AI advice from 1,000 real prompts submitted to GPT-5.2, GPT-5.6, and Gemini 3 Flash.
- AI advice produced sizable saving buffers for nearly everyone over 30, but handled shocks like job loss poorly and let portfolios drift instead of rebalancing them.
- Advice from prompts written by men, the more financially literate, or prior AI users generated about 5% more wealth near retirement; women and less literate users ended up with roughly $50,000 (4%) less wealth at 60, and first-time AI users about $100,000 (6%) less.
- About two-thirds of the gender gap came from men and women phrasing prompts differently; the rest came from the model changing its advice when a prompt was simply labeled as coming from a woman.
- Vanguard and iShares products appeared in 6% and 3.4% of LLM responses respectively, even though fewer than 0.4% of prompts named either company by name.
Why it matters
Half of Americans already say they use AI for financial advice, according to Choukhmane, but until this study there was little evidence on what that advice actually contains or whether following it helps. This is the first attempt to measure the real-world financial impact of AI advice at scale, by simulating what happens to a person's wealth over an entire working life if they follow it.
Who it affects
Anyone using a chatbot for money questions, and disproportionately women, people with lower financial literacy, and people trying AI advice for the first time, who all ended up with measurably less wealth at retirement age under the study's simulation. It also affects financial product providers: the study found LLMs recommending Vanguard and iShares products far more often than users mentioned them, which could shift how firms compete for visibility.
How to use it
The researchers found that detailed, structured prompts, ones that state age, job status, income, savings, life expectancy assumptions, and current tax and Social Security rules, produced better advice than the short, informal questions most people ask. Choukhmane suggested using AI first to build financial understanding rather than simply following its output, and treating it as a complement to a human advisor rather than a replacement, useful for implementing advice in real time between periodic meetings.
How solid is it
The study is grounded in a purpose-built lifecycle model of income, employment, investing, and taxes used as a benchmark for good financial decisions, tested against 1,000 real prompts from adults plus a set of researcher-written academic prompts, then simulated across ages 22 to 89. The paper, "AI Financial Advice: Supply, Demand, and Life Cycle Implications," by Taha Choukhmane, Weidong Lin, and Matthew Akuzawa of MIT Sloan and Tim de Silva of Stanford Graduate School of Business, won the Swiss Finance Institute's Outstanding Paper Award 2026.
Risks and caveats
The AI advice was weaker on judgment calls that fall outside simple rules of thumb, such as adjusting to an income shock or actively rebalancing a portfolio, even under the best, most structured prompts tested. Choukhmane cautioned that there are still no accepted benchmarks for how financial advice should vary by demographics, so it remains unclear how much of the observed gender and literacy gap reflects reasonable tailoring versus bias absorbed from training data.
“We were somewhat surprised by how good the advice was. Especially when you read the kind of questions people asked, it was not a given that the advice would line up with what academics think are good financial principles.”
— Taha Choukhmane, MIT Sloan School of Management