A Visual Guide to RNNs, CNNs, Transformers, and Calibration, with Hands-On Experiments on Accuracy and Efficiency

186 points•Anon84•1 day ago•9 comments•

9 comments

malsheabout 8 hours ago
Sebastian is the author of two excellent books related to LLMs

Build a Large Language Model (From Scratch): https://sebastianraschka.com/llms-from-scratch/

Build a Reasoning Model (From Scratch): https://sebastianraschka.com/reasoning-from-scratch/

Topfiabout 13 hours ago
Solid assessment, very much what I assumed after announcement.

The way quite a lot of brains fell out, some unreflectively quoting how this could get us to AGI, system 1, “no hallucinating”, etc, while others ignored the breadth and new data vs existing classifiers and saw no possible upside, was revealing. The game demos were especially harmful, was told repeatedly that Jev must have near instant visual input support, as few to none of the flashy Doom, Minecraft, etc. showcases explained this was using game state.

Hype really is the worst aspect of this industry.

nzoschkeabout 21 hours ago
Great article. This matches my feelings:

> Like ChatGPT in 2022 was exciting because it was a general-purpose chat model that could generate all kinds of texts, one of the reasons the tech community is excited about Jev is that it is the ChatGPT moment for classification, where it can cheaply classify all kinds of text inputs without having to fine-tune a custom classifier for each task.

We've been comparing strategies for classifying email and Jev is looking promising. I compared some strategies here: https://housecat.com/blog/classifying-email

kevinwangabout 9 hours ago
Really nice explanation of how Jev is and isn't special, something I was struggling to understand before this.
jacek-123about 15 hours ago
Thats very nice, I always explain LLMs to ppl starting from old-school language models and then just replacing the predictor from bag-of-words to transformers to etc.) I think its a cool framing
alansaberabout 13 hours ago
Agreed as well, this is my preferred framing, it only makes sense to get into semantic vectors once you realise the very real limits of discrete vectors and bag of words.

Read the full thread on Hacker News →

Related stories