flash

29 stories and discussions about flash, aggregated from every source we track.

1.

Qwen Studio offers comprehensive functionality spanning chatbot, image and video understanding, image generation, document processing, web search integration, tool utilization, and artifacts.

338 points•jjcm•13 days ago•135 comments•
2.

Gemini 3.8 Flash-Lite TTS and Gemini 3.8 Flash TTS are our most expressive audio models yet.

330 points•swolpers•7 days ago•150 comments•
3.

A foundation model for mobile, wearables, robots, smart home, automotive and microcontrollers. One 8-29 MB binary that beats models 10x its size on mobile tool calls and matches 2-3x bigger models on extraction.

236 points•HenryNdubuaku•13 days ago•92 comments•
4.
113 points•flxflx•4 days ago•49 comments•
5.

You know that Maya Angelou quote that says "Never make someone a priority when all you are to them is an option?" If Flash were a person following that tenet, then it now has to drop Safari from its dwindling list of…

15 points•calvin•over 10 years ago•1 comment
7.

Apple is using a 128GB NAND flash die in Apple Watch Series 12 and Ultra 4 models, claims a Chinese hardware researcher, despite both devices officially offering only 64GB of storage. It's unclear why Apple would…

4 points•bookofjoe•9 days ago•0 comments•
8.

Chinese battery brand Sunwoda has just revealed a new EV battery charging system that can add 100 km (~62 miles)...

3 points•cisc•10 days ago•0 comments•
9.
2 points•OsamaJaber•about 18 hours ago•0 comments•
10.

Analysis of Xiaomi's MiMo-V2.6-Flash and comparison to other AI models across key metrics including quality, price, performance (tokens per second & time to first token), context window & more.

2 points•AnodicElegy•3 days ago•1 comment•
11.

前情提要 每次看到 Gemini 出新功能,我第一個念頭都是「能不能接進我的 LINE Bot」。 9/22 的 Gemini API release notes 寫著...

2 points•evanlin•4 days ago•1 comment
12.

DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm - antirez/ds4

2 points•thombles•5 days ago•0 comments
13.
2 points•worldofmatthew•5 days ago•0 comments•
14.
2 points•mfiguiere•6 days ago•0 comments•
15.

In the Stage 2 report we shipped 20 feature tickets on Fizzy and promised to explore benchmarking the agents all on max-effort. Now we have run it: every model on the board, same tickets, reasoning turned all the way…

2 points•doppp•9 days ago•0 comments•
16.

The fastest way to run Qwen3.8-Flash-Next on Strix Halo (gfx1151) - peonist-ai/halogen-flash-server

1 points•rdslw•about 8 hours ago•0 comments•
17.

TensorFold's new inference engine delivers decode speeds over 62 tokens per second on a single stream and 119 across five, with prefill at 2,500 tokens per second and a 256k context window. It outperforms prior vLLM…

1 points•soltanov•1 day ago•0 comments•
18.
1 points•artur_makly•2 days ago•0 comments•
19.

Analysis of Xiaomi's MiMo-V2.6-Flash and comparison to other AI models across key metrics including quality, price, performance (tokens per second & time to first token), context window & more.

1 points•theanonymousone•3 days ago•0 comments•
20.
1 points•shenli3514•3 days ago•0 comments•
21.

Telemetry audit, post-mortem, and failure-mode analysis of a 474K LOC codebase built with Antigravity + Gemini Flash (102.8B tokens). - Janson79jc/antigravity-100b-telemetry-audit

1 points•janson79jc•3 days ago•0 comments•
22.

Gemini 3.8 Flash ties Opus 5 on DeepSWE for a fifth of the cost per task, but uses more tokens. Flash Cyber is gated; Muse Spark is cheap if Meta trains on you.

1 points•axrisi•3 days ago•2 comments
24.

OpenAI- and Anthropic-compatible inference for coding agents. Call DeepSeek-V4.1-Flash and DeepSeek-V4-Flash with zero data retention.

1 points•handfuloflight•6 days ago•0 comments•
25.

We’re on a journey to advance and democratize artificial intelligence through open source and open science.

1 points•verdverm•7 days ago•0 comments•
26.

Ant Group has released their finance-focused flash model Ling-3.0-flash-Fin

1 points•gmays•8 days ago•0 comments•
27.

Flagship performance · Full modality · Built for professional workflows | Team Token Plan · Batch Inference (Batch API) now powered by V2.6 models.

1 points•smokeeaasd•8 days ago•0 comments•
28.

Same model, same prompt — multiple times the tokens per second, with near-zero rate limits. Inference built for long-running headless agents, so the runs that used to queue now finish on time.

1 points•Hiteshjain118•9 days ago•3 comments•
29.

A 66-minute local Pi coding session on a 64 GB M2 Ultra: 106 tool calls, a 131K context window, automatic compaction, and a working macOS OTelux install.

1 points•b1tank•9 days ago•0 comments•

Related topics