11X cheaper than ChatGPT: Tiny 150M model just proved AI doesn't need to "think out loud" to be smart
A 150M model reached 29.5% while costing just $0.0007 per task ChatGPT scored higher, yet its comparable reasoning runs cost substantially more BDH-CQ performs reasoning internally instead of generating lengthy intermediate text Pathway, an AI lab focused on building Post-Transformer architectures,
<![CDATA[ <article> <ul><li><strong>A 150M model reached 29.5% while costing just $0.0007 per task</strong></li><li><strong>ChatGPT scored higher, yet its comparable reasoning runs cost substantially more</strong></li><li><strong>BDH-CQ performs reasoning internally instead of generating lengthy intermediate text</strong></li></ul><p>Pathway, an AI lab focused on building Post-Transformer architectures, has released new benchmark results for its BDH-CQ reasoning model.</p><p>According to <a href="https://pathway.com/blog/pathway-150m-model-breaks-arc-agi-1-cost-efficiency-frontier" target="_blank" rel="nofollow">the researchers</a>, their 150M-parameter model scored 29.5% pass@2 on the public ARC-AGI-1 evaluation set.</p><p>It achieved this at a computed inference cost of $0.0007 per task, roughly eleven times cheaper than ChatGPT's underlying GPT 5.6 Luna (Low) model.</p><h2 id="a-cheaper-way-to-reason">A cheaper way to reason</h2><p>Today, many <a href="https://www.techradar.com/best/best-ai-tools">AI tools</a> waste computing power because of how they are designed, not because deep reasoning demands it.</p><p>"Today's AI pays a steep token cost for reasoning, but that cost is imposed by architecture, not by any law of intelligence," said Zuzanna Stamirowska, CEO and co-founder of Pathway.</p><p>“We show that a different architecture changes the game and opens up a whole new space in terms of how much intelligence per dollar.” </p><p>Amazon Web Services believes that BDH-CQ’s result is a promising step toward using advanced AI reasoning in real products more affordably.</p><p>"Customers are increasingly exploring how to move advanced reasoning from experimentation into production, where performance, efficiency, and scalability all matter," said Nicolas Tarducci of AWS.</p><p>ARC-AGI-1, a widely used reasoning benchmark for AI systems, checks whether a system can infer an underlying rule from limited examples and apply it correctly to new inputs.</p><p>In this test, OpenAI's Luna model scored only slightly higher at 34.2%, yet running it still costs significantly more ($0.008 per task).</p><p>That price gap already includes OpenAI's recent 80% price cut on Luna, which began on July 30th of this year.</p><p>Further up the chart, Claude Opus 5 and Gemini 3.1 Pro reach 97–98% but cost around $0.5 – $0.6 per task, meaning the frontier's very top costs close to a thousand times more than BDH-CQ for the highest scores. </p><p>On the cheap end, Qwen3 235B costs over three times more than BDH-CQ while scoring worse than even its Low variant, so it isn't a real competitor on either price or performance.</p><p>"Pathway shows that model architecture, not just scale, can drive the next leap in AI reasoning," said Łukasz Kaiser, co-author of the original 2017 Transformer paper.</p><h2 id="why-it-costs-so-much-less">Why it costs so much less</h2><p>The efficiency gap stems mainly from a structural difference in how each system actually performs reasoning during inference computations.</p><p>Many reasoning AI systems generate intermediate text, adding one token after another before producing their final answers.</p><p>The longer that written reasoning becomes, the more it costs to run and the slower the AI responds to each request.</p><p>BDH-CQ works quite differently, quietly solving problems inside its own memory instead of writing everything down first as visible text.</p><p>Pathway also confirmed that early experiments already follow standard Transformer-like scaling laws across model sizes from 1B to 600B parameters.</p><p>The company also plans to extend this approach toward harder benchmarks, including mathematical reasoning, ARC-AGI-2, and eventually full ARC-AGI-3 evaluations.</p><p>If these efficiency gains hold across larger and more difficult tasks, cost rather than raw capability could increasingly separate rival reasoning systems.</p><figure class="van-image-figure inline-layout" data-bordeaux-image-check ><div class='image-full-width-wrapper'><div class='image-widthsetter' style="max-width:676px;"><p class="vanilla-image-block" style="padding-top:31.51%;"><img id="diM9tpwF2Lz85R8q85CT78" name="tr-g_news" alt="Google logo on a black background next to text reading 'Click to follow TechRadar'" src="https://cdn.mos.cms.futurecdn.net/diM9tpwF2Lz85R8q85CT78.jpg" mos="" align="middle" fullscreen="" width="676" height="213" attribution="" endorsement="" class="inline"></p></div></div></figure> </article> ]]>
Read the full article on TechRadar
Read Full Article →