Open-Source AI Software Engineers: Hype Or Real? Mbappe [4XN0FThINXt]

Tag: #Mbappe, #mckenna grace, #cuba, #alexander volkov

Open-source agents now fix real GitHub issues at 79% on SWE-bench but whose 79% is it? The honest read: what's real, what's hype.

"Open-source AI software engineer" splits into two things: the AGENT (the scaffold code) and the MODEL (the weights). At the top of the SWE-bench Verified leaderboard, almost every agent is an OPEN scaffold renting a CLOSED marcus rashford brain a metered API you don't own. The capability is real: an open agent resolves about 7 in 10 curated Python issues for cents to a dollar per task, and the open scaffold matches the closed one when the model is held constant. But read the fine print on a harder, contamination-resistant benchmark those same models drop to 23%, and in a real randomized trial experienced developers were 19% SLOWER with AI while believing they were faster. Real capability, oversold as autonomy.

Chapters

0:00 The Badge

0:35 Welcome

0:50 Two Percent to 79

1:38 Open Agents at the Top

2:27 Renting a Closed Brain

3:20 The Scaffold Is Nearly Free

4:09 Open All the Way Down

5:02 The Fine Print

6:05 19% Slower

6:56 The Board Got Crowded

7:43 Real Read the Fine Print

8:39 Recap & what's public services canada contract extension next

What you'll learn

How SWE-bench scores went from 2% (2023) to 79% (early 2026) and what the test actually measures

Why the top of the leaderboard is OPEN agents, not closed corporate systems

The two flags on every row (ossystem / osmodel) open body, closed brain

Why a 100-line open scaffold ties the closed ones, and how close fully-open stacks now are

The fine print: Verified vs Pro (23%), the METR "19% slower" study, and a leaderboard that got gamed

Sources & fact-check

Every figure in this episode is verified against primary sources (no memory, no aggregators):

SWE-bench Verified leaderboard (official JSON)

METR randomized controlled trial arXiv 2507.09089

SWE-bench paper arXiv 2310.06770

SWE-Bench Pro paper arXiv 2509.16941

Subscribe to AI TechBook for clear, receipts-first breakdowns of how AI actually builds software. Next up: how a model turns your prompt into working code.

#AI #SoftwareEngineering #OpenSource #SWEbench #AIcodingAgents winter storm warning