Meta publicly launched a new version of Muse Spark on Thursday, a multimodal AI model designed for agentic coding that aims to compete with similar products offered by OpenAI and Anthropic. Spark 1.1, ...
Most widely cited AI coding benchmarks, including the original SWE-bench, were built primarily around Python repositories, meaning headline performance results may not accurately predict how coding ...
Get ready for another Meta app. Meta has a new social AI app called Pocket, Business Insider has learned. Meta describes the Pocket as a platform to "create, share, and discover gizmos with friends." ...
The move comes as Musk works to improve the coding capabilities of his Grok family of AI models. Musk recently combined his AI company, xAI, with SpaceX to form a mega-sized firm, which he took public ...
Security leaders from Datadog, Jamf, and ASOS weigh in on the visibility crisis quietly unfolding as AI puts code-writing capabilities in every employee's hands. "I spent the weekend burning through ...
AI coding agent startup Niteshift has raised a $7 million seed round led by Greylock’s Jerry Chen. That’s a modest sum by AI standards, but the startup, founded by two former early Datadog engineers, ...
Anthropic is releasing Claude Fable 5 for general users. Fable 5 uses Mythos-class power with safety controls. Pricing is about twice that of Claude Opus 4.8. Anthropic has announced a defanged ...
For decades, corporate hiring has favored candidates who could present a flawless résumé and deliver highly structured answers to interview questions. Today, generative AI is making it easier for ...
After scathing accusations of skimping on due diligence, as well as other feedback to my article on trying to use an ‘AI coding assistant’ for the first time, the only rational, academic response is ...
The latest flare-up in the debate over AI-assisted coding did not come from a new model release or a benchmark result. It came from a single line of text buried inside a software update. Earlier this ...
The controversy over vibe coding reached a new high this week after a developer added hidden instructions to his open source Java testing app to sabotage projects performed by AI coding agents. The ...
DeepSWE, created by DataCurve offers a benchmark for assessing AI coding models by focusing on real-world programming challenges rather than synthetic test cases. According to Matthew Berman, one of ...
Some results have been hidden because they may be inaccessible to you
Show inaccessible results