GPT-6 Launches With Near-Frontier Intelligence at Half the Price
GPT-6 Sol and Luna halve OpenAI's API prices yet score near rival flagships on coding, while flagship Astra is OpenAI's first Critical-rated cyber model.
Category desk
Technical deep dives, reproducible tests, and tool evaluations.
GPT-6 Sol and Luna halve OpenAI's API prices yet score near rival flagships on coding, while flagship Astra is OpenAI's first Critical-rated cyber model.
Alibaba confirmed Qwen 4 is in training at Apsara 2026, with Qwen 4.5 and Qwen 5 targeting 5-10T parameters and the Zhenwu V900 chip due Q1 2027.
Xiaomi's MiMo-V2.6 clears CyberGym at 94.0 under an MIT licence, but ExploitBench at 47.9 shows the audit-to-exploit gap is still wide open.
In agentic workflows, prefill latency and prefix caching matter more than raw tok/s. We benchmark SGLang RadixAttention vs vLLM PagedAttention and APC.
Choosing a $6,000 local AI workstation: Apple Silicon Unified Memory (128GB-256GB) vs Dual RTX 5090 (64GB). Sizing 70B/236B MoE models, TTFT, and TCO.
Can the RTX 5090 run 70B models? We break down 32GB VRAM sizing, Blackwell NVFP4 vs FP8 benchmarks, KV cache math, and vLLM vs SGLang configurations.
Claude Code CLI burns 400k to 2.5M tokens per task. We analyze subscription tiers (Pro, Max 5x, Max 20x) vs direct API pricing with prompt caching.
Cursor Ultra, Claude Max 20x, and ChatGPT Pro all charge $200/month. We stress-test real rate limits, usage pools, and token ROI during 8-hour sprint days.