Leave your feedback Share Copy URL https://gographicsoutput.com/video/ociXa8dgdUF.html Email Facebook Twitter LinkedIn Pinterest Tumblr Share on Facebook Share on Twitter Open-Source AI Software Engineers: Hype Or Real? Mbappe [4XN0FThINXt] Health Updated on August 05, 2026 EDT — Published on August 05, 2026 EDT Tag: #Mbappe, #mckenna grace, #cuba, #alexander volkovOpen-source agents now fix real GitHub issues at 79% on SWE-bench but whose 79% is it? The honest read: what's real, what's hype."Open-source AI software engineer" splits into two things: the AGENT (the scaffold code) and the MODEL (the weights). At the top of the SWE-bench Verified leaderboard, almost every agent is an OPEN scaffold renting a CLOSED marcus rashford brain a metered API you don't own. The capability is real: an open agent resolves about 7 in 10 curated Python issues for cents to a dollar per task, and the open scaffold matches the closed one when the model is held constant. But read the fine print on a harder, contamination-resistant benchmark those same models drop to 23%, and in a real randomized trial experienced developers were 19% SLOWER with AI while believing they were faster. Real capability, oversold as autonomy. Chapters0:00 The Badge0:35 Welcome0:50 Two Percent to 791:38 Open Agents at the Top2:27 Renting a Closed Brain3:20 The Scaffold Is Nearly Free4:09 Open All the Way Down5:02 The Fine Print6:05 19% Slower6:56 The Board Got Crowded7:43 Real Read the Fine Print8:39 Recap & what's public services canada contract extension next What you'll learn How SWE-bench scores went from 2% (2023) to 79% (early 2026) and what the test actually measures Why the top of the leaderboard is OPEN agents, not closed corporate systems The two flags on every row (ossystem / osmodel) open body, closed brain Why a 100-line open scaffold ties the closed ones, and how close fully-open stacks now are The fine print: Verified vs Pro (23%), the METR "19% slower" study, and a leaderboard that got gamed Sources & fact-checkEvery figure in this episode is verified against primary sources (no memory, no aggregators): SWE-bench Verified leaderboard (official JSON) METR randomized controlled trial arXiv 2507.09089 SWE-bench paper arXiv 2310.06770 SWE-Bench Pro paper arXiv 2509.16941 Subscribe to AI TechBook for clear, receipts-first breakdowns of how AI actually builds software. Next up: how a model turns your prompt into working code.#AI #SoftwareEngineering #OpenSource #SWEbench #AIcodingAgents winter storm warning
Tag: #Mbappe, #mckenna grace, #cuba, #alexander volkovOpen-source agents now fix real GitHub issues at 79% on SWE-bench but whose 79% is it? The honest read: what's real, what's hype."Open-source AI software engineer" splits into two things: the AGENT (the scaffold code) and the MODEL (the weights). At the top of the SWE-bench Verified leaderboard, almost every agent is an OPEN scaffold renting a CLOSED marcus rashford brain a metered API you don't own. The capability is real: an open agent resolves about 7 in 10 curated Python issues for cents to a dollar per task, and the open scaffold matches the closed one when the model is held constant. But read the fine print on a harder, contamination-resistant benchmark those same models drop to 23%, and in a real randomized trial experienced developers were 19% SLOWER with AI while believing they were faster. Real capability, oversold as autonomy. Chapters0:00 The Badge0:35 Welcome0:50 Two Percent to 791:38 Open Agents at the Top2:27 Renting a Closed Brain3:20 The Scaffold Is Nearly Free4:09 Open All the Way Down5:02 The Fine Print6:05 19% Slower6:56 The Board Got Crowded7:43 Real Read the Fine Print8:39 Recap & what's public services canada contract extension next What you'll learn How SWE-bench scores went from 2% (2023) to 79% (early 2026) and what the test actually measures Why the top of the leaderboard is OPEN agents, not closed corporate systems The two flags on every row (ossystem / osmodel) open body, closed brain Why a 100-line open scaffold ties the closed ones, and how close fully-open stacks now are The fine print: Verified vs Pro (23%), the METR "19% slower" study, and a leaderboard that got gamed Sources & fact-checkEvery figure in this episode is verified against primary sources (no memory, no aggregators): SWE-bench Verified leaderboard (official JSON) METR randomized controlled trial arXiv 2507.09089 SWE-bench paper arXiv 2310.06770 SWE-Bench Pro paper arXiv 2509.16941 Subscribe to AI TechBook for clear, receipts-first breakdowns of how AI actually builds software. Next up: how a model turns your prompt into working code.#AI #SoftwareEngineering #OpenSource #SWEbench #AIcodingAgents winter storm warning