serving
9 stories and discussions about serving, aggregated from every source we track.
A step by step deployment of Gemma 4 E2B to a single AMD Instinct MI300X on AMD Developer Cloud, driven by Python MCP tools, and the throughput a 191.7 GiB card returns for its hourly rate.
Compare serving configurations using measured GPU timings, then let an agent implement and validate promising optimizations in real serving frameworks.
GPUs are expensive, yet inference-serving GPU clusters remain heavily underutilized. To improve utilization, state-of-the-art systems adopt GPU multiplexing. However, optimizing solely for utilization can…
What the Databricks and ClickHouse benchmark discussion raised about serving high-concurrency analytics directly from external tables—and how StarTree scaled it past 125K QPS.
Biology is often seen as an experimental science, with theory serving to generalize patterns from empirical data. But history reveals that theory plays a more generative role in discovery itself.
Starting with Feast v0.65, a new tighter integration between Feast and ScyllaDB brings vector search support and better performance to the online store. This post covers what changed, why teams outgrow their existing…
A dataset hub for LLM serving research. Request traces, agent workloads, and GPU telemetry from Harvard MadSys and collaborators.