serve

15 stories and discussions about serve, aggregated from every source we track.

1.

Google's QAT Gemma 4 E2B keeps its embedding tables in bf16, and on a Tesla T4 they are most of the model. Packing them to int4 on the grid QAT trained them onto cuts model loading from 6.33 to 2.86 GiB, with every greedy test output token-identical, and raises vLLM's output throughput 11-37% over Google's own W4A16 export.

8 points•xbill•about 15 hours ago•0 comments
2.

Pope Leo XIV, who has made AI a central component of his pontificate, warned against the dangers of the technology at UNESCO. #EuropeNews

7 points•jethronethro•5 days ago•1 comment•
3.

Learn how we optimized performance and efficiency serving the workload that is changing software engineering forever — and you can too.

4 points•birdculture•5 days ago•0 comments•
5.
6.

Learn how we optimized performance and efficiency serving the workload that is changing software engineering forever — and you can too.

2 points•birdculture•about 3 hours ago•0 comments•
7.

Learn how we optimized performance and efficiency serving the workload that is changing software engineering forever — and you can too.

2 points•birdculture•6 days ago•0 comments•
8.

System One judgments (noul/choice/score) from any LLM in one prefill — a Rust, Jev-compatible /v1/systemone engine - yijunyu/jev-rs

2 points•yijunyu•8 days ago•1 comment•
9.

Model providers now advertise context windows large enough to hold a codebase or a stack of contracts...

2 points•digitalocean_staff•9 days ago•0 comments
10.

We couldn't provide detailed pricing to the CLD suppliers."

1 points•rbanffy•2 days ago•0 comments•
11.

Learn how we optimized performance and efficiency serving the workload that is changing software engineering forever — and you can too.

1 points•charles_irl•7 days ago•0 comments•
12.

Serve analytics agents with context that stays in sync as schemas and business logic change. Every update is tested, reviewed, and approved.

1 points•matthieu_bl•7 days ago•0 comments•
13.

Learn how we optimized performance and efficiency serving the workload that is changing software engineering forever — and you can too.

1 points•kkm•7 days ago•0 comments•
14.

Contribute to taeold/djev-run development by creating an account on GitHub.

1 points•_js•7 days ago•0 comments•
15.

Single logit inference runtime for all LLM models. Contribute to rreinold/jev-serve development by creating an account on GitHub.

1 points•rreinold2•8 days ago•0 comments•

Related topics