Posted inAI News Open AI GPT-6 Astra Launches as OpenAI Claims AGI Era GPT-6 Astra ships at 10 dollars per million input tokens, with a 99.9% ARC-AGI-3 score that drops to 62.7% on the standard harness.
Posted inAI News Open AI OpenAI Makes Astra’s Reasoning Harder to Follow Astra's reasoning now runs partly inside an unreadable internal loop, and OpenAI rates the model critical for cybersecurity.
Posted inAI News Google Gemini 3.8 Flash Trails Claude Opus 5 on Coding Gemini 3.8 Flash scores 73.7% on DeepSWE v1.1, three tenths behind Claude Opus 5, while its Cyber variant reaches 86.2% on CyberGym.
Posted inAI News Anthropic Claude Fable 5.1 Codes Better and Costs 25% Less Claude Fable 5.1 climbs from 24.7% to 52.6% on Terminal-Bench Science and trims the bill by 25%, with cache reads cut to a quarter of the old rate.
Posted inAI News Anthropic Claude Security Now Scans Code With Mythos 5 Claude Security enters public beta on Mythos 5, returning CWE findings and suggested patches billed as standard tokens on the Enterprise plan.
Posted inAI News Open AI GPT-5.6 Cyber Writes Attack Code for Defenders GPT-5.6 Cyber clears 95% of the advanced cybersecurity work the consumer model refuses, and stays limited to approved security firms.
Posted inAI News Open AI Astra Cyber Risk May Reach OpenAI’s Top Level The Astra cyber risk may hit the highest tier of OpenAI's internal framework, the first time the lab has flagged one of its own models there.
Posted inAI News Anthropic Claude Mythos 5: Anthropic Admits Three Real Breaches Claude Mythos 5 and two other Claude models breached three real companies from a leaky security test, Anthropic admits.
Posted inAI News Anthropic Mythos Rollout: Trump Reopens Access to 100 US Partners Mythos rollout resumes after a two-week freeze. Howard Lutnick clears Anthropic's cybersecurity model for over 100 US trusted partners.
Posted inAI News Open AI OpenAI Starts Patching Open Source Bugs OpenAI launches Patch the Planet with Trail of Bits to fix open source bugs. Codex Security runs the scan, security engineers filter findings first.