Measured tests of cloud AI products
Frontier on Cloud is a set of communities, one per cloud, that test cloud AI products at the frontier and publish the measurements. Every number on this site comes from a public repository at a named commit, with the raw data committed next to the summary. Each README has a Reproduce section that reruns the test in three commands: copy the example environment file, install, run. Where a setup needs more than that, the test page says so. The rules are on the Standard page.
Communities
| Cloud | Community | Status |
|---|---|---|
| Google Cloud | r/FrontierOnGCP | Active. All four tests below. |
| AWS | r/FrontierOnAWS | Coming |
| Azure | r/FrontierOnAzure | Coming |
Tests
Newest first. Each page has the setup, the exact numbers, the caveats and the commit the numbers come from.
-
A pending tool call across a lost connection and session resumption
Measured: what happens to a pending booking when the Live API WebSocket is lost, and what session resumption recovers. Result: after a silent loss every resume was refused with close code
1011while the server held the old connection (450 of 450 attempts over 15 minutes with real packet loss); a client-side fallback to a new session answered correctly in 9 of 9 runs. -
LiveKit agents-js#2615 on gemini-3.8-live: reproduction and patch check
Measured: whether a tool result is lost after a content-free
generationComplete, as the issue reports, and whether the patch proposed in the issue fixes it. Result: unpatched, the result was lost in 2 of 2 runs where the model spoke before calling the tool and delivered in 3 of 3 where it called first; patched, delivered in 5 of 5. -
A client-side commit guard for in-flight tool calls, before and after
Measured: a reference guard (hold the commit on speech onset, dedupe by business key, send the model a status note, handle abandoned BLOCKING calls) on the stop test's speech scenarios. Result: a booking committed after the user's stop in 12 of 12 runs with the guard off and in 0 of 12 with it on.
-
What "stop" does to an in-flight tool call on Gemini 3.8 Live
Measured: whether the Live API cancels a pending
book_slotcall when the user says "stop", with text and speech input. Result: notoolCallCancellationin any of the 41 sessions; withbehavior: BLOCKINGthe model re-issued the call and the slot was booked twice in 3 of 3 text runs.