gemma

24 stories and discussions about gemma, aggregated from every source we track.

1.

The behind-the-scenes story of how we built DinoDesk AI—a privacy-first, LEGO dino companion robot powered by a hybrid Local Gemma 4 and Cloud Gemini architecture.

44 points•bebechien•15 days ago•7 comments
2.

Serving Gemma 4 E2B q4_0 through llama.cpp on one laptop, twice: CPU-only and on a 2021-era 4 GB GTX 1650 Ti. Same GGUF, same binary, same prompts, one flag apart. The card takes decode by 4.3x, and needs only 1598 MiB to do it.

15 points•xbill•14 days ago•6 comments
3.

Building the smallest Compute Engine VM that serves Gemma 4 E2B on one Tesla T4, installing the driver and vLLM after boot, and a walkthrough of every option in the shell script that starts, checks and queries the server.

13 points•xbill•8 days ago•0 comments
4.

Gemma 4's 4-bit builds on SageMaker's smallest GPU, an NVIDIA T4, against the L4: a Turing patch for vLLM, the host image the CUDA 13 container needs, speed, memory, answers and cost per token.

12 points•xbill•about 7 hours ago•1 comment
5.

A step by step deployment of Gemma 4 E2B to a single AMD Instinct MI300X on AMD Developer Cloud, driven by Python MCP tools, and the throughput a 191.7 GiB card returns for its hourly rate.

11 points•xbill•13 days ago•5 comments
6.

Gemma 4 E2B, E4B, 12B and a 26B-A4B fp8 build read by their label probabilities with vLLM on one TPU v6e chip, checked against the same read on an NVIDIA L4 and against Jev 1.13.0's published results. What fits one chip, how to read labels when vLLM on TPU returns only the top 32 log-probabilities, speed, cost, and why no 31B loads today.

10 points•xbill•6 days ago•0 comments
7.

Gemma 4 E2B q4_0 served by llama.cpp on one laptop, CPU-only and on a GTX 1650 Ti, rebuilt on CUDA 13.4 and re-measured in CPU, GPU, GPU, CPU order with a temperature gate. The card takes decode by 4.14x, and run order moves the answer by about 2%.

10 points•xbill•7 days ago•2 comments
8.

Step-by-step: running Google's quantization-aware-trained Gemma 4 E2B on a 10th-gen Core i7 laptop with a 4 GB GTX 1650 Ti — why bf16 and int8 cannot fit, why the QAT GGUF does with room to spare, and managing it with an MCP server.

10 points•xbill•20 days ago•3 comments
9.

Plain Gemma 4 26B read by its label probabilities against DiffusionGemma's one-step read, both as community 4-bit (AWQ) builds on one EC2 L4, on 1,200 labelled examples and on the 3,880-record public suite where Jev 1.13.0 has published results. Pre-registered, with accuracy, calibration, calibration after 0 to 150 labels, latency and cost.

8 points•xbill•7 days ago•0 comments
10.

How I replaced a cloud LLM with a fully local one — same agent, same tools, zero inference cost.

8 points•mazlum_tosun•14 days ago•0 comments
11.

Step by step deployment of Gemma 4 E2B to a Cloud Run NVIDIA L4 GPU with vLLM, managed by a Python MCP server migrated to the MCP SDK 2.x.

8 points•xbill•20 days ago•1 comment
12.

Google's QAT Gemma 4 E2B keeps its embedding tables in bf16, and on a Tesla T4 they are most of the model. Packing them to int4 on the grid QAT trained them onto cuts model loading from 6.33 to 2.86 GiB, with every greedy test output token-identical, and raises vLLM's output throughput 11-37% over Google's own W4A16 export.

7 points•xbill•about 12 hours ago•0 comments
13.

A short background on SageMaker real-time endpoints, then a measured comparison of Gemma 4 E2B's QAT w4a16 checkpoint against the full-size bf16 release on the same NVIDIA L4 endpoint: decode speed, parallel throughput, answers and cost.

7 points•xbill•5 days ago•0 comments
14.

Step by step deployment of Gemma 4 E2B to a SageMaker real-time endpoint on one NVIDIA L4 with the AWS vLLM container, driven by the aws CLI and managed by a Python MCP server from Claude Code or Gemini CLI.

7 points•xbill•5 days ago•0 comments
15.

A step by step deployment of Gemma 4 E2B with vLLM on a single Tesla T4 attached to a Compute Engine VM, and a measured comparison of the QAT w4a16 checkpoint against the bf16 reference on the same card.

7 points•xbill•12 days ago•1 comment
16.

Local finite-choice decisions with Gemma — yes/no probabilities, named choices and ordered scores in Go.

2 points•rcarmo•7 days ago•0 comments•
17.

A gallery-local explainer for per-layer embeddings in Gemma 4 E2B and E4B.

2 points•Bluestein•9 days ago•0 comments•
18.

The same Gemma 4 build, vLLM version and GPU served from a SageMaker endpoint and from a plain EC2 instance, on a T4 and an L4: identical decode and answers, a different call path, and what the managed endpoint's 1.40x buys.

1 points•xbill•about 6 hours ago•0 comments
19.

Gemma 4's 4-bit builds on SageMaker's smallest GPU, an NVIDIA T4, against the L4: a Turing patch for vLLM, the host image the CUDA 13 container needs, speed, memory, answers and cost per token.

1 points•xbill•about 7 hours ago•0 comments
20.

Plain Gemma 4 26B read by its label probabilities against DiffusionGemma's one-step read, both as community 4-bit (AWQ) builds on one EC2 L4, on 1,200 labelled examples and on the 3,880-record public suite where Jev 1.13.0 has published results. Pre-registered, with accuracy, calibration, calibration after 0 to 150 labels, latency and cost.

1 points•xbill•7 days ago•0 comments
21.

Gemma 3 4B as a Jev-style System One decision function: one forward pass, choice logits, softmax. Native Rust + Candle. No generate(). - zozo123/gemma-to-jev

1 points•zozo123-IB-IL2•8 days ago•0 comments•
22.

Repacking Gemma 4's QAT weights five ways and serving each on the same SageMaker NVIDIA L4 endpoint: int4 linears, int4 embeddings and lm_head, FP8 and int8, across E2B, E4B, 12B, 26B A4B and 31B.

0 points•xbill•about 12 hours ago•0 comments
23.

A short background on SageMaker real-time endpoints, then a measured comparison of Gemma 4 E2B's QAT w4a16 checkpoint against the full-size bf16 release on the same NVIDIA L4 endpoint: decode speed, parallel throughput, answers and cost.

0 points•xbill•5 days ago•0 comments
24.

Step by step deployment of Gemma 4 E2B to a SageMaker real-time endpoint on one NVIDIA L4 with the AWS vLLM container, driven by the aws CLI and managed by a Python MCP server from Claude Code or Gemini CLI.

0 points•xbill•5 days ago•0 comments

Related topics