embeddings

6 stories and discussions about embeddings, aggregated from every source we track.

1.

Google's QAT Gemma 4 E2B keeps its embedding tables in bf16, and on a Tesla T4 they are most of the model. Packing them to int4 on the grid QAT trained them onto cuts model loading from 6.33 to 2.86 GiB, with every greedy test output token-identical, and raises vLLM's output throughput 11-37% over Google's own W4A16 export.

7 points•xbill•about 12 hours ago•0 comments
2.

Hello, I'm Rijul, and I'm building LiveReview — a blast-radius aware AI code review built for your...

7 points•rijultp•5 days ago•0 comments
3.
3 points•softwaredoug•2 days ago•0 comments•
4.

Banking77: 94.25% accuracy with a 642 KB logistic classifier (embeddings + sklearn) - banking77_gist.py

2 points•nico•4 days ago•0 comments•
5.

A gallery-local explainer for per-layer embeddings in Gemma 4 E2B and E4B.

2 points•Bluestein•9 days ago•0 comments•
6.

Repacking Gemma 4's QAT weights five ways and serving each on the same SageMaker NVIDIA L4 endpoint: int4 linears, int4 embeddings and lm_head, FP8 and int8, across E2B, E4B, 12B, 26B A4B and 31B.

0 points•xbill•about 12 hours ago•0 comments

Related topics