• OpenAI scientist Noam Brown: The true upper limit of AI may not be able to measure

    As large language models gradually tackle complex tasks such as reasoning, automated research, and cybersecurity, traditional methods of evaluating models are facing new challenges.For a long time, the release of models has been accompanied by a report of results consisting of various benchmark tests in areas such as mathematics, programming, scientific question answering, network security, and knowledge reasoning, which are then compared horizontally with the previous generation of models.

    OpenAI scientist Noam Brown: The true upper limit of AI may not be able to measure
  • Claude secretly becomes stupid while doing AI research, and Anthropic is besieged by the research community

    Claude Fable 5 is the main focus in the AI field today; this “mythical” model performs exceptionally well and has attracted tremendous attention.Andrej Karpathy described it as “very exciting” and an “evolutionary step that deserves a major version upgrade,” on par with the improvements brought by Claude 4.5 last November. In the SWE-bench Pro programming benchmark, Fable 5 achieved a score of 80.3%, which is 11 percentage points higher than Opus 4.8. With a Ruby codebase consisting of 50 million lines of code, Fable 5 completed the entire codebase migration in just one day; if the same task were assigned to a human team, it would take more than two months.

    Claude secretly becomes stupid while doing AI research, and Anthropic is besieged by the research community
  • GPT-5.6 The first batch of actual measurements are here! Precise sniper on Mythos

    Just now, Anthropic released its “big killer” – a product that has been in development for two months.Claude Fable 5andMythos 5It's like dropping a bomb.Apply direct pressure on OpenAI now.

    GPT-5.6 The first batch of actual measurements are here! Precise sniper on Mythos

No More