What changed

  1. Mistral: Mistral Small 3.2 24Bcontext window 131072 → 256000source

  2. Qwen: Qwen3.8 27Bcontext window 1048576 → 1000000source

  3. Z.ai: GLM 4.7 Flashcontext window 202752 → 200000source

  4. Qwen: Qwen3.8 27Bcontext window 1000000 → 1048576source

Observed mechanically from public sources; every line carries the source it was read from. The last 30 days as a table · RSS

Impossible → Routine

Capabilities that were research results, and the date each became something anyone could buy. Both ends dated, both ends sourced.

A forecast that was learned, not simulated

Forecasting global weather with a learned model instead of simulated atmospheric physics.

Impossible

GraphCast, a research model, beats the industry gold-standard physics forecast on more than 90% of tested variables.

10-day forecast in under a minute on one TPU machine

source

2 years, 9 months

Routine

WeatherNext 3 becomes the new leader on Brightband's Operational WeatherBench, a live third-party comparison of AI and physics global medium-range models, edging out its predecessor WeatherNext 2.

lowest 2m-temperature error on 26 of the last 30 days in August

source

Real work on a real desktop

Completing multi-step tasks on a real computer — opening applications, clicking through interfaces, finishing the job.

Impossible

The OSWorld benchmark arrives and the best model completes 12.24% of its 369 real-computer tasks.

12.24% on OSWorld, the original 369-task set

source

2 years, 3 months

Routine

Qwen3.8-Max posts 86.1% on OSWorld-Verified, the repaired 2025 revision of the benchmark, with competing agents from two other labs already above 83%.

86.1% on OSWorld-Verified, the 2025 revision

source

All 27 dated pairs