157 points•1vuio0pswjnm7•10 days ago•89 comments•

89 comments

ehe78qhe10 days ago
Article is just a vague summary of https://www.saturnos.com/report/artificial-authority

Anecdotally, current models seem to be decent at general personal finance principles - certainly better than the majority of personal finance education that people get exposed to unless they seek it out and read a variety of books and sources. But I wouldn't trust them with direct decision making with actual money due to the training lag time on current tax policy, etc.

NitpickLawyer10 days ago
> certainly better than the majority of personal finance education that people get exposed to unless they seek it out and read a variety of books and sources.

Also, models are now good enough that you can give them chapters from "authoritative" books, and they'll integrate that and come up with better answers even if their "vanilla" answers were average. And they'll tailor stuff to your particular situation. It's funny that the "agentic" stuff is only used in coding mostly, while it can and does work in other fields as well.

As always, you kinda need to check it (at least spot check) but all in all I'd agree it's better than the average stuff you used to find with a quick google search.

legostormtroopr10 days ago
So if you know a book that has the information you need, you just need to upload it into the model to get the right answer.

Isn't that a bit circular - if you already know the authoritative source, why ask a model?

FearNotDaniel10 days ago
Important to note is that what is being measured here is the ability of the models not of the chat tools themselves, which combine model completions with other tools that the models can call upon. The mainstream labs already know this about models, it's no secret, and in fact training materials from e.g. Anthropic are at pains to point out that users, or analysts designing workflows, have the reponsibility to ensure the correct tools are used and that human verification takes place at appropriate stages depending on the risk/consequences of the task at hand.

Of course a language-completion model with a training cutoff date won't have up-to-date information on tax rules or the ability to carry out correct numerical calculations, but when you combine that with (in Claude terminology) web search and code execution tools invoked by the chat agent, you immediately have much more reliable results.

ehe78qhe10 days ago
I keep finding that the current harnesses, when encountering syntax that was invalid at training time but is now valid due to new language versions or custom extensions, don't correctly figure out why and assume something is wrong with the codebase or toolchain. I would hate to have that happen with my taxes.
htrp10 days ago
I guess the question becomes, how much of this is harness versus model?
AnimalMuppet9 days ago
Well, lag time against current tax law is definitely model.
johnnienaked10 days ago
It has very little to do with lag time on tax policy
bluecalm10 days ago
So I downloaded that report which of course doesn't contain the most relevant information (the questions) but it contains some examples of wrong answers.

I fed the first question to Grok (which they claimed they tested as well) and it answered it correctly in detail.

I repeated it with another one - again correct answer. I then selected the question they said Grok specifically answered incorrectly and it again answered it correctly.

I am sticking with my first intuition: people are terrible at testing tools and probably wanted them to answer incorrectly/not fully (the questions are constructed in a way to make it difficult as well). They also have vested interest in the conclusion (they are financial advisory firm) so there is that to consider.

People reading ft will now think chat boxes are bad at answering financial questions while they are pretty good at it. Zero consequences for spreading fake news for Financial Times there but good for financial advisors I guess.

ModernMech10 days ago
> each LLM was tested 600 times, and in total over 10,000 questions and answers were assessed.

Okay but why do you feel 3 trials say as much as 10,000?

bluecalm10 days ago
>>Okay but why do you feel 3 trials say as much as 10,000?

I don't trust them so I've used 3 examples in incorrect questions/answers they have given and I got correct answers. I spend enough time with LLMs to know that if Grok answered it correctly and in detail then it wouldn't be a problem for GPT or Claude either.

The questions are also constructed in a way that it's easy to answer not fully (which they qualify as wrong). LLMs still answer them correctly and in detail though.

Dwedit10 days ago
LLMs use random numbers, so a single test won't necessarily match someone else's experience.
bluecalm10 days ago
They use random numbers so the answers sounds a bit different but in my experience you won't be able to get a wrong answer to a simple question no matter how many times you try it.
ozgung10 days ago
This will be the same story for every industry again and again. AI is not good in x=Finance because models were not RL trained heavily on x=Finance capabilities yet. This is only because Big Labs have finance benchmarks lower in their priority list. Their first priority was solving programming because that gives the best leverage at this stage. As a side-effect they were able to solve a Millennium Prize problem, since theorem proving was also code.

So finance advisors in the comments section of FT are falling for the classical pitfall. They assume there is something fundamentally wrong with “AI chatbots” that they can’t do finance ever. They mistake the current products in the market for the technology itself. In near future someone will release “Claude x=Finance” and their world will shatter.

otabdeveloper410 days ago
Programming hasn't been solved by LLMs. AI chatbots give wrong answers to programming queries most of the time too.
qgin10 days ago
Not most of the time
high_na_euv10 days ago
Not really

I do find them useful when querying like "how todo xyz in abc"

anfogoat10 days ago
> AI is not good in x=Finance because models were not RL trained heavily on x=Finance capabilities yet.

Not sure what the "financial queries" here amounted to but I find it hard to believe LLMs will ever be to finance what they are to programming. Past a point, information related to the former is gatekeeped behind private institutions with special government granted privileges, while information related to the latter is freely available and open to anyone.

ImaCake10 days ago
Plenty of code is not publically available, and yet that has not stopped them from RLHF'ing good coding bots. You can surely see the parallels to any other field including finance.
lukeify10 days ago
Given most financial advisors tend to vend out suboptimal advice and steer customers in favour of products they receive a kickback for, I'm happy to be accepting of an unbiased LLM that's trained on bogleheads.org.
in_absentia10 days ago
If that's what you want, I'll save you some tokens:

#!/bin/sh

while read question; do echo "Put it into VFIAX"; done

ehe78qhe10 days ago
This is missing a lot of steps like:

- Building an emergency fund

- Budgeting and tracking where your money goes

- Planning and saving for large purchases like cars, homes and life goals

- Optimizing use of tax-advantaged accounts like 401Ks, HSAs, and IRAs

- What to do with ESPPs, RSUs, and options

- How taxes work and how to optimize around them

- Estate planning

koito1710 days ago
There's quite a few implicit assumptions in that.

In my case, I am double-taxed (both Japan and US side) on capital gains. Tax treaties reduce, but not eliminate, the extent of double-taxation.

Many US-based brokers do not allow Americans abroad to purchase mutual funds, so VFIAX is not a choice for me.

Maybe I can go with eMAXIS Slim All Country... Oh, but that is a PFIC under IRS rules and I'd be taxed on unrealized capital gains. So I guess no Japan-equivalents of VT for me. That's fine, I guess I'll just buy VT in my US-based brokerage account; but now I'm in a suboptimal spot with respect to monthly contributions, calculating JPY-denominated income tax on dividends, etc.

I even made an implicit assumption when I said "calculating JPY-denominated income tax on dividends". That assumes your tax status is permanent resident. If your tax status is non-permanent resident, then a decent financial advisor will recognize that only the extent of income remitted to Japan gets taxed, so VT distributing at all isn't an issue (until 5 years later). What should you do before the 5 year threshold is hit? etc. etc.

But yes, if you're born in America and plan to stay within the same state for the rest of your life, then a 100% automated setup that simply deposits $1,000/mo into VFIAX is probably fine. (But keep in mind, to most non-Americans, VOO is not really diversified compared to funds like VT).

Otherwise, there is value in consulting someone (or something) that regularly handles taxes and financial planning.

IshKebab10 days ago
Some people do have complex financial situations. It's not as simple as that.

For example in the UK (and maybe US?) you get tax relief for money you put into your pensions, but there's a limit of £60k/year. Unless you earn a lot (which I do, yeay) when that limit is tapered. Except that you can also use up to 3 years of previously unused allowance. But you have to use this year's first.

Also interest is taxed, but you can put up to £20k/year into an ISA which isn't. And if you still want to avoid some tax you have kids ISA's and even pensions!

Then there are also startup investment schemes that save you some tax. Those seem to be not worth it, but you get the idea - it can be complicated. Especially if you are near one of the many tax/benefit thresholds.

The marginal tax rate in the UK bounces all over the place - it's even technically possible for it to be over 100%!

chasil10 days ago
While I also practice Bogle's approach from The Little Book of Common Sense Investing, even with this baseline there are some subtleties.

-VFIAX is currently $707/share. Fidelity's FXAIX does not have to be purchased in increments of a share price, and this fund's expenses are lower.

-There are versions of the S&P 500 for taxable accounts that minimize capital gains.

-Vanguard has a total-market index, VTSAX, that is mentioned in the book.

-Vanguard also has a non-U.S. total market fund, VTIAX, that avoid the current CAPE problems of the U.S. market.

Claude is very familiar with Bogle's approach, likely because the pirated book was part of the training set.

schnitzelstoat10 days ago
It depends though - I'm not in the US and in my country ETFs get taxed continually whereas mutual funds don't get taxed until you sell.

In general, invest in low-cost index funds is pretty solid advice everywhere. In different countries you might use slightly different instruments due to tax advantages (like the ISA in the UK etc.)

derwiki10 days ago
That’s why it’s important to get a Fiduciary Financial Advisor who does not get kickbacks based on fund choices.
AnimalMuppet9 days ago
How long before at least one model is steering customers toward products that it receives a kickback for?
lukeify9 days ago
Yeah that's a good point. Bring on open weight models.

Read the full thread on Hacker News →

Related stories