Close Menu

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    Countertop Materials and Everyday Kitchen Wear

    August 29, 2026

    20 Wild Self-Diagnoses That Were Actually Right

    August 29, 2026

    Chinese automakers are following Tesla’s bet that robots are the next big profit machine

    August 29, 2026
    Facebook X (Twitter) Instagram
    Trending
    • Countertop Materials and Everyday Kitchen Wear
    • 20 Wild Self-Diagnoses That Were Actually Right
    • Chinese automakers are following Tesla’s bet that robots are the next big profit machine
    • My Most Contrarian Opinion Right Now
    • Announcer Icon Sasses Trump During NFL Broadcast — And It’s Quite The Hit
    • Twitch Is Down | Lifehacker
    • Busy Philipps Found Out She Had A Brain Tumor Just Days After James Van Der Beek Died
    • What Is WebMCP? How to Prepare Your Website to Serve AI Agents
    Facebook X (Twitter)
    SBM Global News
    Demo
    • Home
    • Top Stories
      • Politics
    • Business
      • Small Business
      • Marketing
    • Finance
      • Investment
    • Technology

      Chinese automakers are following Tesla’s bet that robots are the next big profit machine

      August 29, 2026
      Read More

      Service Robot Co. – Company Profile

      August 28, 2026
      Read More

      Google’s new Fitbit Air brings Pokémon Sleep to your wrist

      August 28, 2026
      Read More

      The LegalTech AI Company Seeing Enormous Traction

      August 27, 2026
      Read More

      India’s Ringg gets backing from Peak XV as it pushes voice AI past the phone call

      August 26, 2026
      Read More
    • Lifestyle
      • Travel
    • Feel Good
    • Get In Touch
    SBM Global News
    Demo
    Home»Technology»Did xAI lie about Grok 3’s benchmarks?
    Technology

    Did xAI lie about Grok 3’s benchmarks?

    By Staff WriterFebruary 23, 20253 Mins Read
    Facebook Twitter LinkedIn Reddit Email
    #image_title
    Share
    Facebook Twitter LinkedIn Pinterest Email

    Debates over AI benchmarks — and how they’re reported by AI labs — are spilling out into public view.

    This week, an OpenAI employee accused Elon Musk’s AI company, xAI, of publishing misleading benchmark results for its latest AI model, Grok 3. One of the co-founders of xAI, Igor Babushkin, insisted that the company was in the right.

    The truth lies somewhere in between.

    In a post on xAI’s blog, the company published a graph showing Grok 3’s performance on AIME 2025, a collection of challenging math questions from a recent invitational mathematics exam. Some experts have questioned AIME’s validity as an AI benchmark. Nevertheless, AIME 2025 and older versions of the test are commonly used to probe a model’s math ability.

    xAI’s graph showed two variants of Grok 3, Grok 3 Reasoning Beta and Grok 3 mini Reasoning, beating OpenAI’s best-performing available model, o3-mini-high, on AIME 2025. But OpenAI employees on X were quick to point out that xAI’s graph didn’t include o3-mini-high’s AIME 2025 score at “cons@64.”

    What is cons@64, you might ask? Well, it’s short for “consensus@64,” and it basically gives a model 64 tries to answer each problem in a benchmark and takes the answers generated most frequently as the final answers. As you can imagine, cons@64 tends to boost models’ benchmark scores quite a bit, and omitting it from a graph might make it appear as though one model surpasses another when in reality, that’s isn’t the case.

    Grok 3 Reasoning Beta and Grok 3 mini Reasoning’s scores for AIME 2025 at “@1” — meaning the first score the models got on the benchmark — fall below o3-mini-high’s score. Grok 3 Reasoning Beta also trails ever-so-slightly behind OpenAI’s o1 model set to “medium” computing. Yet xAI is advertising Grok 3 as the “world’s smartest AI.”

    Babushkin argued on X that OpenAI has published similarly misleading benchmark charts in the past — albeit charts comparing the performance of its own models. A more neutral party in the debate put together a more “accurate” graph showing nearly every model’s performance at cons@64:

    Hilarious how some people see my plot as attack on OpenAI and others as attack on Grok while in reality it’s DeepSeek propaganda
    (I actually believe Grok looks good there, and openAI’s TTC chicanery behind o3-mini-*high*-pass@”””1″”” deserves more scrutiny.) https://t.co/dJqlJpcJh8 pic.twitter.com/3WH8FOUfic

    — Teortaxes▶️ (DeepSeek 推特🐋铁粉 2023 – ∞) (@teortaxesTex) February 20, 2025

    But as AI researcher Nathan Lambert pointed out in a post, perhaps the most important metric remains a mystery: the computational (and monetary) cost it took for each model to achieve its best score. That just goes to show how little most AI benchmarks communicate about models’ limitations — and their strengths.



    View original article here

    Share. Facebook Twitter LinkedIn Email Reddit
    Previous ArticleTrapped on Boat for a Decade
    Next Article The 7 Principles of Conversion-Centered Landing Page Design

    Related Posts

    Chinese automakers are following Tesla’s bet that robots are the next big profit machine

    August 29, 2026
    Read More

    Service Robot Co. – Company Profile

    August 28, 2026
    Read More

    Google’s new Fitbit Air brings Pokémon Sleep to your wrist

    August 28, 2026
    Read More
    Add A Comment

    Leave A Reply Cancel Reply

    Demo
    Top Posts

    Former FBI, CIA Head Has ‘Serious Concerns’ With Trump Cabinet Picks

    December 28, 2024435

    Emirates to operate next-gen A350 on the third daily service to Cape Town

    January 14, 2026256

    AAVE Price Prediction: Target $215-225 by Mid-January 2025 as Technical Indicators Signal Bullish Momentum

    December 15, 2025240

    Ventive Hospitality Joins Green Fins: Strong ESG Lift

    February 17, 2026211
    Don't Miss
    Lifestyle

    Countertop Materials and Everyday Kitchen Wear

    By Staff WriterAugust 29, 20268 Mins Read

    What this covers The Four Properties That Matter Quartz and Quartzite Are Not the Same…

    Read More

    20 Wild Self-Diagnoses That Were Actually Right

    August 29, 2026

    Chinese automakers are following Tesla’s bet that robots are the next big profit machine

    August 29, 2026

    My Most Contrarian Opinion Right Now

    August 29, 2026
    Stay In Touch
    • Facebook
    • Twitter
    Demo
    About Us

    Small Business Minder brings together business and related news from around the world in one place. Follow us for all the business news you'll need.

    Facebook X (Twitter)
    Our Picks

    Countertop Materials and Everyday Kitchen Wear

    August 29, 2026

    20 Wild Self-Diagnoses That Were Actually Right

    August 29, 2026
    Most Popular

    Former FBI, CIA Head Has ‘Serious Concerns’ With Trump Cabinet Picks

    December 28, 2024435

    Emirates to operate next-gen A350 on the third daily service to Cape Town

    January 14, 2026256
    © 2026 Small Business Minder
    • Home
    • Get In Touch

    Type above and press Enter to search. Press Esc to cancel.

    Ad Blocker Enabled!
    Ad Blocker Enabled!
    Our website is made possible by displaying online advertisements to our visitors. To get the most from our site, please disable your Ad Blocker.