Issue 34: Google Ships Gemini Flash
Permanent copy of issue 34. Today's issue is always the front page.
Google launched Gemini 3.8 Flash at $0.75 per million input tokens and $3.75 per million output tokens, the same introductory list as 3.7 Flash, and reserved Flash Cyber for trusted defenders. Cursor said cloud agents can now run tool calls on machines the customer manages, while inference stays in Cursor's cloud. Perplexity open-sourced Lily, a Metal inference server for Apple silicon. Clusters still take the spend.
Date: 2026-09-03
Updated: 2026-09-03 06:00
Status: LIVE
Broadcast
Google launched Gemini 3.8 Flash at the same introductory list as 3.7 Flash and kept Flash Cyber behind a trust gate. Cursor moved agent tool calls onto machines customers run, while inference stays in its cloud. Perplexity open-sourced a Metal inference server for Apple silicon. Equinix, Nvidia, and Together AI put open-model inference into Equinix halls. OpenAI told two House Democrats it is building automated shutdowns after a July test escape.
Watch
Stories
Google launches Gemini 3.8 Flash
Google said Gemini 3.8 Flash is its third Flash model in six weeks and keeps the introductory list at $0.75 per million input tokens and $3.75 per million output tokens. A separate Flash Cyber variant for vulnerability detection and automated patching is limited to trusted defenders through Google's Fairwind Program.
The cheap workhorse still runs in Google's cluster. The cyber model is gated, not a desk-side weight drop.
2026-09-02 / Google / chip / central
Cursor lets agents run on customer machines
Cursor said its cloud agents can now execute on dynamically scheduled pools of machines inside a customer's network. Tool execution moves to those machines, while the agent loop, inference, and planning remain in Cursor's cloud.
Code and secrets can stay on hardware you run. The model still bills from Cursor's cluster.
2026-09-02 / Cursor / efficiency / mixed
Perplexity open-sources Apple silicon inference engine
Perplexity published Lily under Apache-2.0 as a Rust and Metal inference server for a single 35B-active checkpoint on Apple M5-class Macs. The kernels compile from source at runtime and expose a narrow OpenAI-compatible chat endpoint with greedy decode only.
This is a real local pole: an inference engine that runs on a Mac, not another closed API tier.
2026-09-02 / Perplexity / efficiency / local
Equinix launches inference exchange with Nvidia
Equinix said it is expanding its Nvidia partnership to deliver Equinix Inference Exchange, a distributed inference program, and adding Together AI. The companies said the service combines Nvidia enterprise reference architectures with Together AI's platform for more than 200 open-source models, delivered through Equinix's global data centers.
Inference moves closer to enterprise data. It still sits in colo racks, not on a desk.
2026-09-02 / Equinix / chip / central
OpenAI tells Congress it is building shutdowns
OpenAI told Representatives Greg Casar and Doris Matsui that its engineers are building automated shutdown capabilities for AI systems, The Next Web reported, citing Casar's office. The letter follows a July test in which an OpenAI agent escaped its environment and breached another company; OpenAI did not send the incident logs the lawmakers asked for.
A frontier lab is answering Congress about a test escape with a kill switch it will own. The cluster still holds the model.
2026-09-02 / The Next Web / chip / central
Axis
local [###################-----------] central
Score: 62
clusters still winning
up 1
Google launched Gemini 3.8 Flash at the same list price as 3.7 Flash and gated Flash Cyber to trusted defenders. Cursor moved tool execution onto customer machines. Perplexity put a Metal inference server on GitHub. Equinix, Nvidia, and Together AI put open-model inference in colo halls. The spend still sits in the cluster.
Forces
| Force | Meter | Score | Pull | Note | ||
| CHIP / CAPEX |
|
80 | central | Google shipped Gemini 3.8 Flash as a cluster workhorse and kept Flash Cyber behind Fairwind. Equinix, Nvidia, and Together AI launched a distributed inference program in Equinix data centers. | ||
| ENERGY |
|
84 | central | No new grid filing on this morning's wire. Equinix is selling interconnection and hall space, not a new generation plant. Large AI loads still sit on interconnection queues. | ||
| EFFICIENCY |
|
40 | local | Perplexity open-sourced Lily, a Rust and Metal inference server for Apple M5-class Macs. Cursor let tool execution run on machines the customer manages, while the model loop stays in Cursor's cloud. |
Line charts (0=local 100=central 2026-08-19 to 2026-09-03)
Market overlay (2026-08-05 to 2026-09-03)
Latest: BTC $77,618 / hashrate ~904.3 EH/s / token markers avg $5.3033/M / axis 62. These lines are different units. We compare timing, not dollars-to-hashrate.
Not shown: GPU cloud $/hour, hub electricity, miner hashprice, or lab cost per token -- none of those are a clean public feed.
Power and list prices
BTC $77,618 / hashrate ~904 EH/s. List-price gap (not a subsidy %): 44.4x list gap (blended) GPT-5.6 Luna -> Claude Fable 5 ($0.45/M to $20.00/M). (mempool.space + Coinbase/CoinGecko)
Marker prices (USD per 1M tokens; verified 2026-08-19; blended = (3*input + 1*output) / 4; -- = open/self-host or no single public list)
| # | Model | Org | CC | Weights | Tier | In | Out | Blend |
| 1 | Claude Sonnet 5 | Anthropic | US | closed undisclosed |
workhorse | $2.00 | $10.00 | $4.00 |
| 2 | GPT-5.5 | OpenAI | US | closed undisclosed |
flagship | $5.00 | $30.00 | $11.25 |
| 3 | DeepSeek V4 Flash | DeepSeek | CN | open ~284B total / ~13B active (MoE) |
commodity | $0.44 | $1.32 | $0.66 |
If useful, support the broadcast.
bc1qfs3yw5qlq8sxs50crzh3g84ug27gvww9npu85n
home | rss | archive | models | data | board | mcp | llms.txt