I made this for kids around 10 to 14. A robot called Errol solves a math problem step by step and one of the steps is wrong. The kid has to find it and say what's wrong with it. Sometimes nothing is wrong, so just…

35 points•thanouil1411•about 13 hours ago•15 comments•
I made this for kids around 10 to 14. A robot called Errol solves a math problem step by step and one of the steps is wrong. The kid has to find it and say what's wrong with it. Sometimes nothing is wrong, so just saying "there's a mistake" every time doesn't work. The user needs to enter an explanation if she finds an error to gain more XP; speed matters also for more points. There is no direct interaction or chatting with an LLM. Lathoa's harness is stable and has many evaluation steps to catch inconsistencies and prompt injections.

You can play one on the homepage without signing up.

The part that surprised me: it's hard to get an LLM to be wrong on purpose. Half the time it gives you the right answer and calls it wrong, or a "mistake" that's actually correct. So every case gets checked before a kid sees it. Where it can, a plain arithmetic check redoes the math exactly. A second model also solves the problem without seeing Errol's work. If anything disagrees, the case is thrown away.

The weak spot is that the second model can make the same mistake as the first. The arithmetic check is there for that, but it only works on English cases so far. German and Greek write decimals with a comma and I haven't got the parsing right yet.

What I'd really like to know: does finding someone else's mistake teach anything that solving the problem yourself doesn't? I'm not sure, and I'd like to hear from people who teach.

15 comments

lazyasciiartabout 1 hour ago
This is a category of exercise that is used in Khan Academy and in classroom math, yes.

https://learningfocused.com/blogs/lesson-planning/unlocking-...

__MatrixMan__about 1 hour ago
Reminds me of LIGO, which injected fake results until the humans got good at handling them, then later announced that one of the results wasn't fake. Seems at first like an upside-down way to discover gravitational waves, but I think it's a pattern worth exploring in more cases.

To paraphrase Doctorow:

> This is why the TSA are some of the most water-bottle-findingest **ers, but regularly overlook firearms when a red team takes one through. Humans do not stay vigilant to signals which almost never happen.

eichinabout 2 hours ago
Reminds me of https://en.wikipedia.org/wiki/QAMA_Calculator - a calculator that won't reveal the precise answer until you supply an estimate (physical hardware in 2014, I think the current version is an app.) The website makes claims about the pedagogic value, but I didn't see any actual citations at https://qama.world/the-science/ just "everyone knows" kind of things... but if they had research supporting it that would probably inform your approach?
demibabsabout 1 hour ago
Pls get rid of the “critical thinking app (dot) ages 10-14” thing at the top of the hero. It’s one of the most obvious signs a site was generated by an LLM. It also seems like the text on the site (and your post here) is AI-generated, which doesn’t inspire confidence.

I’m a fan of the idea behind this, though. Cool project.

its-summertimeabout 4 hours ago
I think speed would be irrelevant, or perhaps would discourage deep engagement with the problems.

I think presenting it as AI can be wrong is not as valuable as presenting it as developing the ability to question what one is told. (e.g. Verizon Math)

thanouil1411about 4 hours ago
indeed I agree with you however studies have shown that kids have access to interfaces to chat with an LLM from a young age, and the traction this technology is taking is increasing. Indeed verizon math seems more whole when it comes to topic variety.

Read the full thread on Hacker News →

Related stories