I finally built Mathy, a little project I’ve been thinking about for a couple years. It’s free, and no account is required if you want to try it on iOS or web. I use Math Academy daily…
37 comments
I did a little deep dive into this kind of stuff over the summer. I haven't written it up yet, but I'll share some highlights below.
I realized Anki was optimized for the opposite of the problems I actually had. Anki's optimized to minimize study time. It does this by maximizing difficulty. For me, this means maximizing pain!
I'm not low on time; I'm low on willpower. I need it to be fun and easy! Otherwise, I quit and memory goes to zero. Gwern implies I am typical in this regard; most people who try SRS don't stick with it, even if they want to. So "Minimize Pain" seems like a worthy endeavour here.
So, by necessity, I made the reps easy. By making the reps easy, I realized I had accidentally become much more fluent. [0]
(The "automaticity" you mention, comes from practicing more frequently.) In other words, working hard (suffering more) was giving me worse results. Imagine! (I had a similar experience with complex motor skills, but that's a separate post!)
At some point I realized I could replace my entire flashcard app with a lightly modded Anki: increase target retrievability (to increase frequency), and use latency as the score (easy with a tiny addon, and/or custom card type).
As far as the math goes, I got stuck on generating the cards (past arithmetic). How did you do that? :)
(Also, I'd love to see the knowledge graph! I've been trying to reverse engineer the old one from Khan Academy...)
--
[0] The ancients knew this principle well...
https://en.wikipedia.org/wiki/Precision_teaching
Speaking of fluency, I think "overlearning" is a misnomer. Our standards are too low and we are chronically "underlearning". Similarly, our "Mastery Learning" is actually Temporary Basic Competence...
This is relevant because you can't necessarily apply Very Frequent Reps to your entire knowledge base. There's just not enough time to maintain fluency in everything.
But occasionally, you will want to be fluent in something, which you were once fluent in. Then, you simply shift it into Tier A, spam reps, and rapidly regain what was lost -- for as long as it's relevant to you.
To answer your question, I'm a product guy, so it was easier for me to start with the user app/mobile UX I wanted + constraints rather than starting with the cards.
Then I created a problem view in the dev version of the app that let me see every type of problem to see how it rendered, how input worked, etc.
Then with that plus the system around it (details below) gave me higher confidence in generating the 40,000+ problems across all the topics. It still has room for improvement, but happy with how it came out so far.
So, for the problems, they were programmatically generated, but NOT LLM-generated/trusted, so the process would give me the confidence:
1. Define the problem families explicitly. Each has bounded inupts, known mathematical rule, answer type, constraints & presentation rules.
2. A deterministic compiler generated the problems. So, given the same source definitions/versions itd produces the same corpus. so there's no AI inventing random questions at runtime .
3. Correct answers are computed from the underlying math, not from the rendered text. Integer/rational problems use exact arithmetic. More complex symbolic cases use validation recipes, incl. offline SymPy if needed.
4. LaTeX is just presentation (tried other way, didn't work). Internally the problem is represented as structured math, then laTeX/MathJax derives from that structure, so it doesnt generate a LaTeX string and cross fingers its interpreted correctly.
5. Generation/verification are separate steps. Every generated record has to pass mathematical, domain, schema, answer format & presentation checks before it can enter the corpus. Symbolic cases can be independently verified, for example by differentiating a proposed antiderivative.
6. Whole families are tested, not just samples. Finite spaces can be exhaustively enumerated. Larger spaces get property, boundary, invariant, and regression tests. It also runs corpus-wide audits for malformed questions, duplicate IDs, invalid answers, broken rendering, unreachable answer forms, etc
7. The shipped artifact is tied back to what was verified. Versions/content digests bind source definitions, generated problem, validation result & runtime representation together. If something changes, it has to be revalidated rather than reusing old results.
So, starting out I assumed doing it programmatically would be liability. But rather confidence comes BECAUSE it's programmatic. For a large class of problems deterministic generators + exact/exhaustive validation was easier to audit than tens of thousands of hand written questions. So could leverage AI for all that, just had to define the rules/review the system.
I'm very happy with the outcome, but in terms of process it far exceeded my expectations in what I learned about approaching problems like this.
I'm a very heavy AI coding using (I burn hundreds of billions of tokens a year!) so this was a fun way to validate my approach and test some new ones to build high quality experiences, particularly on mobile which is more of a taste thing. The decade plus of product experience really came in handy in guiding the AI here, which took a lot of iteration on the experience and trying different things. It was a blast, and I was generally surprised at how fast it came together. Thank you again.
>1. Define the problem families explicitly. Each has bounded inupts, known mathematical rule, answer type, constraints & presentation rules.
Yeah, this is the "knowledge graph", right? (Or does each node in the knowledge graph have multiple problem families?)
How did you come up with the graph itself? The curriculum.
And did you say you're shipping SymPy on the mobile client?
Very cool, thanks!
One small piece of feedback: on my Pixel, the Android keyboard obscures the green button. So when I'm answering a numerical question, I use your number pad to enter the digits, and then use the enter key on my Android keyboard. This feels weird, but I guess with regular use it wouldn't.
- It turns out good math questions have a ton of forms! Multiple choice, fill in the blank equations, word problems, diagram labeling, etc, and a good math book or online tool varies the images and forms used so students don't get bored. - AI just isn't great at generating these forms directly, so I ended up having to create these heavy question form adapters to try to get the output of AI to work.
At the time I did it, about six months ago, the LLM I could afford to use for the content just wasn't good enough, and I was having a lot of issues with accuracy, adherence to the curriculum item, and a lack of variability in the numbers and words chosen. Though writing this comment actually gives me a few new ideas...
I shelved it but I'm tempted to try again seeing this and knowing how much models have improved.
Programmer of 27 years, doing math by other names. I continue to be frustrated by how bad math-as-a-programming-language is. This tool did nothing to improve that feeling.
The goal of the product is to give practice and help build automaticity in topics you already know, which makes learning harder stuff (e.g. through Math Academy) easier later to raise your ceiling.
So, I'm not sure if maybe you're less familiar with the material (below the threshold) and you should take a course focused on teaching the material OR that the material in Mathy is too hard for the selected topic. Because if you're familiar, it should be helping you rather than frustrating you.
It's an offline app so I don't capture that kind of data. But you make a good point, and at a minimum for each topic I can have it start with easier problems then increasingly get harder as you do better.
I'll think more about that today and maybe get a fix out once I have some time.
Thanks again for the feedback!
Read the full thread on Hacker News →
Related stories
- Hacker News · 1 points · 2 days ago
- Hacker News · 421 points · 6 days ago
- Hacker News · 2 points · about 10 hours ago
- Hacker News · 1 points · 1 day ago
- Show HN: Mathy – Build Math Automaticitymathy.gameHacker News · 1 points · 9 days ago
- Hacker News · 3 points · 3 days ago