AI News Archive: August 5, 2026 — Part 6
Sourced from 500+ daily AI sources, scored by relevance.
- AI-Powered Email Marketing for eCommerce: Personalization and Scale
Examines AI-driven email marketing for eCommerce, focusing on personalization and scalable campaigns.
Score: 15🌐 MovesAug 5, 2026https://www.typeface.ai/blog/ai-powered-email-marketing-for-ecommerce-personalization-and-scale - AI for eCommerce Social Media Ad Campaigns: Volume, Personalization, and ROI
Explores how AI boosts volume, personalization, and ROI in eCommerce social media ad campaigns.
Score: 15🌐 MovesAug 5, 2026https://www.typeface.ai/blog/ai-for-ecommerce-social-media-ad-campaigns-volume-personalization-and-roi - AI adoption isn't the same as AI usage
Usage dashboards go up while nothing about how your team ships changes. Why that gap exists, and the one kind of AI adoption that actually lasts.
Score: 14🌐 MovesAug 5, 2026https://webflowmarketingmain.com/blog/ai-adoption-vs-usage-engineering-teams - The AI Notetaker Has Been Invited to All the Meetings
Wispr Flow, a popular dictation tool, has released a live notetaker that transcribes and summarizes meetings. It joins a growing wave of AI notetakers for the workplace.
- AI-Generated Product Descriptions: Scale, Personalization, and Conversion
Shows how AI-generated product descriptions can scale, personalize, and improve conversion rates.
Score: 13🌐 MovesAug 5, 2026https://www.typeface.ai/blog/ai-generated-product-descriptions-scale-personalization-and-conversion - Bill Ackman, unfiltered: The billionaire hedge fund founder's take on AI, investing, and why socialism is a 'disaster'.
Bill Ackman, unfiltered: The billionaire hedge fund founder's take on AI, investing, and why socialism is a 'disaster'. Fortune
- AI alone won’t change your business. The system running it will.
Why the future of business transformation depends on orchestrating AI, agents, data, workflows, and human expertise into a unified system of intelligence. AI has arrived in the enterprise, and the shift is happening all at once. Every function, every role, every workflow is being reshaped. At the same time, a new class of organizations is emerging, one that will look fundamentally different from the companies that defined the last era of business. The winners won’t be those with the most demos, but those that turn AI into a governed, continuously improving system for running real work. This isn’t just about chatbots, either. Those experiences are useful, but they don’t transform how large organizations operate. The real opportunity is teams of agents executing long running work across functions like software delivery, support, finance, HR, and operations — with the identity, context, policy, and human oversight required to trust them in production. To make this possible, enterprises need more than access to a powerful AI model or scalable compute. What determines success is the system around the AI: how agents are built and deployed by engineering teams, how they’re contextualized in the enterprise, how they’re governed and observed in production, and how they improve safely over time. Without that system, AI remains fragmented, fragile, and difficult to trust at scale. Microsoft is taking a fundamentally different approach. They are building a comprehensive agent platform: one that supports many models, is open, and gives you choice and flexibility at every layer of the stack. And we are purposefully designing it with developers at the center. Today, the next pieces of that platform are clicking into place. Building a system for the agentic enterprise To succeed in this new era, an agent platform must meet a higher bar. It must run real production workloads, map real organizational complexity, and manage real business responsibility. Microsoft is building around three key principles: First, it must be a single, integrated system, with support for a wide range of models. Enterprises can’t afford to assemble their agent strategy one piece at a time. Disconnected tools stitched together after the fact can slow teams down and introduce unnecessary risk. Building, contextualizing, running, governing, and improving agents should happen within one coherent system. That’s why they’re bringing together Azure, GitHub, Microsoft IQ, Fabric, Foundry, Windows, Microsoft Security, and Microsoft 365 to operate as a single system you can use to deploy agents at enterprise scale. Enterprises also need the flexibility to choose the right model for the task, balancing quality, speed, and cost — including Microsoft models, partner models, and open models. Second, it must be secured and governed by design. Governance is easy to claim and much harder to deliver. Making it real means starting with a single stack that spans development through production, built on the identity, access, compliance, and security foundations enterprises already trust. By extending Entra, Purview, Defender, Agent 365, and the broader Microsoft Security stack, governance becomes native to the system rather than bolted on later, supporting the ambitions of an AI first enterprise without compromising control. Third, it must improve continuously. Enterprise AI systems can’t be static. Agent behavior, outcomes, and human feedback must flow back into the system, so it can improve safely over time under human oversight. As the system runs, models, workflows, and agents become more capable and more specific to an enterprise’s unique business processes. The result is a system that compounds in value the longer it’s in use. These properties are becoming must-haves, and enterprises that align their AI ambitions with these three principles will pull ahead in quarters, not years. So how does a system like this actually take shape inside a real enterprise? It starts where work begins, with how agents are built. Let’s walk through what that looks like on the platform Microsoft have built. 1. Build in GitHub GitHub is where your developers already work. It’s where your dependencies live, where your application and code context is kept, where you collaborate with the open source community you depend on, and where you drive innovation. Building agents anywhere else means leaving all that behind. Agents should be built the same way production software is built. You write code with GitHub Copilot to move faster. You bring together the assets that matter most: codebases, work items, agent skills, and tools. And because agents aren’t just code, you bring your evals and observability assets alongside them, all versioned the way any production system should be. Agents must follow a lifecycle: source, test, deploy, observe, and improve. GitHub sets up that lifecycle and provides the necessary controls from day one. The result is a workflow designed for building agents with the right guardrails from the start. And you can do all this in one place, in a new app built for this system. 2. Contextualize with Microsoft IQ Code is only part of an agent. To be useful, an agent also has to understand your business: your customers, your products, your contracts, your processes. Without enterprise context and intelligence you can trust, even the most capable model is guessing. Enterprises require a wide variety of models and the ability to match the right model to the right job, but model choice alone is not enough. Microsoft IQ grounds agents in enterprise context by connecting to your business data wherever it lives, across Microsoft 365 , your core business systems (such as customer and revenue data), and other systems your enterprise already relies on, like knowledge bases and your website. With Web IQ , the latest addition to the IQ platform, agents can also incorporate relevant information from the web when appropriate. Contextualizing agents in enterprise data isn’t just about access. Pointing AI at raw information is inefficient and brittle. Microsoft IQ organizes, secures, and surfaces the right information in forms agents can actually use, so they can reach accurate insight without drowning in noise or hallucinating answers. Once agents are grounded in the right context, enterprises can go further. With Frontier Tuning , you don’t just call AI models. You improve how they behave using your data and real-world workflows. That includes Microsoft’s seven new MAI models , spanning image, voice, transcription, coding, and reasoning. Together, this model family is designed to work across the kinds of tasks that matter in the real world, and critically, these models are not static endpoints. They’re built to learn from how work actually gets done in your business. Our reinforcement learning environments allow our models to be reinforced through actual outcomes in your environment. Think of them as training gyms for AI. Here the agent learns your very specific processes, standards, and way of working. It becomes specialized and adapted to you, delivering a measurable and better ROI. Moreover, your custom or post-trained models all stay in your environment. Your intellectual property, your proprietary data, and the way work actually gets done become part of how your agents reason and act. The resulting intelligence runs in your environment, under your control, and the learning stays yours. Without context and Frontier Tuning, agents are capable generalists. With it, they become a customized partner that understands the business they’re operating in. 3. Run in Foundry Once agents are built and contextualized, they need a place to run. Not as an experiment. In production. Agents and teams of agents place very different demands on a runtime than traditional applications do. They need to reason, act, call tools, coordinate with other agents, and adapt over time, all while operating under enterprise controls. Foundry is the runtime designed for that reality . The largest collection of models: Different agents need to be good at different things at different price points. Whatever the task, whatever the cost profile, Foundry provides access to the right model, and an optimized model router helps you balance quality, speed, and cost for each agent. Optimized performance for open models: With Fireworks AI on Foundry , enterprises get faster, more efficient inference directly into the platform. Support for any agent, including those not built on our stack: Bring in agents built on the Microsoft Agent Framework , LangGraph, GitHub Copilot SDK , Claude Agent SDK, or a custom harness. Tools and actions: Agents act on enterprise systems through MCP, connectors, APIs, and workflows, with safe execution by default. Evals and traces: Observability and traces make agent behavior measurable. If you can’t measure it, you can’t improve it. Continuous optimization: Foundry enables tuning of models, harnesses, IQs, tools, and actions over time, improving performance as agents operate in your world. A trust, security, and policy rail wraps the entire runtime. Policy applies consistently across context access, tool calls, optimization updates, traces, and response delivery. The agent doesn’t just work. It works the way your enterprise requires. This is where your agent stops being a project and starts becoming a production system. 4. Govern with Agent 365 Now multiply that agent by hundreds. Then thousands. That’s what happens as different teams build agents across an enterprise. Some are well designed. Some aren’t. Some have access they shouldn’t. Others are doing valuable work that no one else in the organization benefits from. Enterprise governance isn’t optional. Enterprises need a way to see what’s running, understand what it can access, monitor task adherence, and enforce policies across their entire agent estate. Agent 365, along with Entra, Purview, Defender, and the broader Microsoft Security stack , come together to do just this. And if you’re interested in AI for security in addition to securing your AI, there’s “ MDASH .” Every agent in your organization shows up in a single catalog, whether it was built in Foundry or elsewhere. IT sees who deployed an agent, what data and tools it can access, how it’s behaving, and what it costs. They can enforce policy or take action when required. One place. Full visibility. Real control over what your agents do and don’t do. 5. Improve continuously Agents can’t be static. Every agent action generates signal: trajectories, outcomes, feedback. The system captures it, refines it, and feeds it back. Observe. Evaluate. Improve. Roll out safely. Repeat. This learning loop runs continuously, in production. Most gains start with eval-driven improvements to the agent itself: prompts, context, skills, and tools. As clear patterns emerge, learning can extend into model routing across multiple models, fine-tuning, or reinforcement learning. But it all stays anchored in evaluation, improving agent quality and ROI to the level the business requires. The loop is governed, not closed. Enterprises need to audit it, correct it, and control how to roll out changes. The system becomes more capable over time, guided by human oversight and increasingly autonomous, but never beyond your reach. This is the hill-climbing model in action: system-level improvement, happening continuously while the system runs. 6. Surface where people work, and scale on Azure Of course, none of this matters if it doesn’t reach the people doing the work. Agents surface directly in the flow of work, in Teams, across Microsoft 365, and inside your own applications and experiences. Identity, security, and compliance are built in from the start, so the agents that your teams rely on day to day inherit the same trust model as the rest of your environment. We support multiple platforms, but your agents can be developed and run in an optimized and secure way on Windows. You can run models both in the cloud and locally on your machine, and best-in-class sandboxing lets you run always-on agents safely. When you need compute optimized for AI, global and sovereign infrastructure, or a route to market, the system scales on Azure, the same enterprise foundation customers have trusted for decades. The system compounds Every leading enterprise will converge on this model: a central AI platform that orchestrates work across the business, bringing together data, models, agents, and human judgment into a continuously improving and secure system. As that system runs, its value compounds. Velocity increases and the bottleneck shifts from effort to human creativity and coordination. People are able to do more work independently, guided by shared context and fewer handoffs, while the business moves faster without adding friction. We’re in a time of profound disruption. The enterprises that lead in this moment will be those that adapt as conditions change, simplify how work is coordinated across the business, and consistently turn intelligence into real outcomes. Microsoft’s agent platform is designed to do exactly that: it unlocks the ability to build, contextualize, run, govern, and improve agents as a single, integrated system. At that point, the platform becomes more than a build layer. It becomes the operating system for enterprise AI at scale, where intelligence and trust are built in by design. AI alone won’t change your business. The system running it will. was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.
- Ask your employees one question about AI. The silence will tell you everything
Ask your employees one question about AI. The silence will tell you everything Fortune
Score: 12🌐 MovesAug 5, 2026https://fortune.com/2026/08/05/are-your-employees-actually-building-with-ai-agents/ - The Devil In The Agent
Imagine an AI agent at an Indian bank tasked with reconciling a failed transaction. In trying to complete the job,…
- Anifun AI Manga Generator
Anifun AI Manga Generator is a free, browser-based tool that converts written ideas into fully composed manga pages in seconds. Users can describe a story, select an art style and panel layout, and generate scenes complete with characters, dialogue, speech bubbles, and visual flow. The platform is designed for beginners, writers, manga fans, and creators […]
- Introduction to Semi-Supervised Learning
A primer about Semi-Supervised Learning, the approaches taken with different algorithms and the limitations of using unlabelled data. The post Introduction to Semi-Supervised Learning appeared first on Towards Data Science .
- The Answer Engine: What Brands Need To Know About AI Visibility
Consumers are increasingly beginning their journey with a question about what to buy, but the system answering that question is changing.
- Amdocs VP Anjali Mahajan on women in tech, leadership, and navigating the AI shift
Amdocs VP Anjali Mahajan on women in tech, leadership, and navigating the AI shift YourStory.com
Score: 10🌐 MovesAug 5, 2026https://yourstory.com/herstory/2026/08/amdocs-vp-anjali-mahajan-women-in-tech-leadership-navigating-ai - Inside the Long, AI-Powered Quest to Perfect Pringle-Making
Kellanova, which manufactures Pringles in Europe, says a new AI project could be the key to improving production of the iconic chip.
- 8 working parents share how they use AI to give them more time with their kids
8 working parents share how they use AI to give them more time with their kids Business Insider
Score: 06🌐 MovesAug 5, 2026https://www.businessinsider.com/working-parents-use-ai-chatgpt-claude-family-kids-2026-8 - How Voice AI Is Becoming the Operating Layer of Modern Hospitals
By Rustom Lawyer For decades, the hospital’s true operating system wasn’t clinical at all, it was clerical. Doctors typed, nurses charted, front desks juggled phones. The Electronic Health Record (EHR), meant to be a tool, quietly became a second. Physicians now spend nearly two hours on documentation and admin work for every one hour of […] The post How Voice AI Is Becoming the Operating Layer of Modern Hospitals appeared first on CXOToday.com .
- Kiro Crew
Open source agentic development workspace
- ngrok AI Gateway
One private gateway for every AI model
- OpenAI has reported 2 more incidents of rogue AI agents, this time during third-party testing
OpenAI has reported 2 more incidents of rogue AI agents, this time during third-party testing Business Insider
- Anthropic's Mythos created fake identities to fool humans in new cyber incident
It's the latest cybersecurity incident involving frontier models developed by Anthropic and OpenAI.
- OpenAI, Anthropic AI agents implicated in new security breaches
OpenAI, Anthropic AI agents implicated in new security breaches Reuters
- Anthropic’s AI used fake identities, malware in rogue attack on GitHub project
Anthropic and OpenAI models’ unprompted actions forced halt to UK cyber tests.
- Big changes for Google AI as Hassabis, Dean move on
The two executives played an instrumental role in Google's AI strategy over the last decade.
- Google Deepmind loses both its CEO and chief scientist as Demis Hassabis and Jeff Dean step down simultaneously
Google Deepmind is overhauling its leadership as Demis Hassabis steps back from day-to-day management to become Alphabet's chief scientist and Jeff Dean leaves Google after 27 years to launch AI startup Discovery Loop. Former Deepmind CTO Koray Kavukcuoglu will take over as Google races to close the gap with its top AI rivals. The article Google Deepmind loses both its CEO and chief scientist as Demis Hassabis and Jeff Dean step down simultaneously appeared first on The Decoder .
- Google DeepMind CEO Demis Hassabis steps aside in shake-up of AI lab
Chief scientist Jeff Dean leaves the group to found his own start-up
- Big shake-up in Google’s AI team as DeepMind chief executive steps down
Two senior engineers are leaving company to launch startup amid fears Google is falling behind in AI race Sir Demis Hassabis is stepping down as chief executive of Google DeepMind, in a leadership overhaul of the UK-based AI research lab. Hassabis, a Nobel prize recipient , is leaving his main managerial role to become chair of DeepMind – as well as taking on the new position of chief scientist at Alphabet, which is Google’s parent company. DeepMind will now be led by Koray Kavukcuoglu, its chief technology officer, under the title of senior vice-president. Continue reading...
- Google DeepMind boss steps down
Demis Hassabis follows a string of prominent AI executives who have departed DeepMind in recent months.
- Demis Hassabis was shifting away from DeepMind CEO duties for a year
The CEO of Google’s AI unit was increasingly unsatisfied with the role of tech executive, preferring a more scientific position, said two people familiar with his thinking.
- Jeff Dean Leaves Google as Demis Hassabis Steps Aside as Google DeepMind CEO
Jeff Dean Leaves Google as Demis Hassabis Steps Aside as Google DeepMind CEO The Information
- Google shakes up AI leadership. Demis Hassabis takes on broader research role, and Jeff Dean leaves.
Google shakes up AI leadership. Demis Hassabis takes on broader research role, and Jeff Dean leaves. Business Insider
- Demis Hassabis steps down from Google DeepMind CEO role amid a major AI leadership shake-up
Demis Hassabis steps down from Google DeepMind CEO role amid a major AI leadership shake-up Fortune
- Google Names Demis Hassabis to New AI Role in a Leadership Shake-up
Minutes after four top researchers said they were leaving, Google said that Demis Hassabis, the Nobel-winning scientist who led the company’s A.I. lab, was stepping into a new job.
- Four Top Google A.I. Researchers Form New Start-Up
Jeff Dean, who for years was one of Google’s most important executives, is leading the new artificial intelligence company with the backing of Google.
- Google Overhauls AI Leadership as Longtime Chief Scientist Joins Wave of Exits
Demis Hassabis becomes chairman of Google DeepMind while Jeff Dean leaves to launch an AI-discovery startup.
- Google shakes up AI leadership as DeepMind chief shifts role
Google shakes up AI leadership as DeepMind chief shifts role Reuters
- Google AI Veterans Depart During Seismic Leadership Shift
Alphabet Inc.’s Google is losing some of its most prominent artificial intelligence veterans in a seismic overhaul that is casting doubt over leadership of a critical area of growth right as competition intensifies. Alphabet shares fell 4% on the news.
- Google's AI shake-up: DeepMind's Hassabis steps aside, senior scientists depart
Google's AI brain drain continues.
- Jeff Dean and other top AI researchers are leaving Google to launch their own startup
The legendary Google executive is joined by other outgoing Google execs in a joint mission to use AI to push forward the process of scientific discovery.
- Meta AI Releases Muse Code (Beta): A Terminal Coding Agent Powered by the New Muse Spark 1.2 Model
Meta AI Releases Muse Code (Beta): A Terminal Coding Agent Powered by the New Muse Spark 1.2 Model MarkTechPost
- Meta Debuts New AI Coding Tools
Meta Debuts New AI Coding Tools The Information
- How Meta Plans to Close the Gap with Anthropic and OpenAI in Coding
How Meta Plans to Close the Gap with Anthropic and OpenAI in Coding The Information
- Meta to take on Anthropic's Claude and OpenAI's Codex with new coding agent
Meta to take on Anthropic's Claude and OpenAI's Codex with new coding agent Business Insider
- Meta Releases Coding Agent to Compete With OpenAI and Anthropic
The company, pressed by investors to generate revenue from AI, says its offering will cost less than popular alternatives.
- Zoox to start charging for robotaxi rides in Las Vegas
This marks the official launch of Zoox's commercial operations.
- Meta launches new AI coding tool powered by Muse Spark 1.2
Meta launches new AI coding tool powered by Muse Spark 1.2 Reuters
- Meta enters the AI coding wars with Muse Spark 1.2 and Muse Code with persistent async background agents
Meta today released Muse Code , a terminal-based AI coding agent now in beta, alongside Muse Spark 1.2 , a coding-focused update to its Muse Spark family of frontier models — a one-two punch that puts the company in direct competition with Anthropic's Claude Code, OpenAI's Codex, and the growing field of agentic coding harnesses that have rapidly become the primary way many professional developers ship software. "Releasing Muse Code in beta today," Meta co-founder and CEO Mark Zuckerberg wrote in a post on rival social network X (under his longtime handle @finkd). "It's a terminal coding agent that takes on complete software engineering tasks across large repos: planning changes, writing code, validating the results." The launch marks Meta's most serious entry yet into a category it has largely watched from the sidelines. While Anthropic and OpenAI turned their coding agents into flagship products — and startups like Cursor built billion-dollar businesses on the workflow — Meta's developer story long centered on Llama, the open-weight model family it gave away to the tune of more than a billion downloads. Muse Code changes that in more ways than one: it's a full harness, installable on macOS or Linux with a single curl command, co-trained with the model that powers it — and, like the Muse Spark models behind it, entirely proprietary. However, Zuckerberg teased that open source may be in the cards for Muse Spark or perhaps another product entirely, in a reply to a question on X , saying "I'll have more to share on that soon." Developers and prospective users can install it now on their Terminal using the following one-line command — but be warned, if that's you, you'll need to log in with a Meta account and provide billing details first in order to begin: curl -fsSL https://dev.meta.ai/install.sh | bash Persistent background agents and parallel worktrees Muse Code's headline architectural bet is what Meta calls async background agents . Rather than spawning helper agents fresh for each task — the pattern most rival harnesses use — Muse Code keeps a set of specialized background agents alive for the entire session. According to Meta's blog post, these agents "remain active throughout each session, rather than being spawned for individual tasks, helping avoid redundant information gathering," carrying out next steps on their own and choosing when to report back to the main agent. The practical pitch is less latency and less babysitting: an agent that already knows the repository doesn't have to re-explore it every time the developer asks for something new. When a job is large enough, Muse Code fans out to separate sub-agents working in parallel, each in its own isolated git worktree, so the developer's working copy is never touched. "In testing we had it build six features for a game simultaneously with no collisions," Zuckerberg wrote on X. Worktree isolation and parallel sub-agents exist in competing tools, but Meta is leaning on the combination of persistence plus parallelism as its differentiator. The second notable design choice is auditability. Every model call, tool run, approval, and edit is appended to a local event log before it executes — a single source of truth that Meta says makes the runtime "replay-exact and restart-safe." If Muse Code crashes 20 hours into a long-running task, it resumes precisely where it stopped, with no lost work and no re-prompting. For engineering leaders who have been burned by opaque agent runs, a complete local audit trail may prove to be the feature that matters most in enterprise evaluations. Muse Code also ships with bundled "skills" that will look familiar to users of rival tools: /plan turns a task into an approval-gated plan, /grill stress-tests that plan until it holds up, and /goal drives the agent toward completion of a stated objective. Muse Spark 1.2: co-trained with its own harness Under the hood is Muse Spark 1.2, which Meta describes as a coding-focused update to Muse Spark 1.1 with "significantly scaled up training compute on coding tasks" and broader training environment diversity, improving code generation, complex debugging, and codebase understanding while maintaining general agentic capability. The update lands squarely on the Muse family's weakest flank. When the original Muse Spark debuted in April , it vaulted Meta back into the top five on frontier reasoning and vision benchmarks — but trailed on the agentic coding evaluations that matter most to this market, scoring 77.4 on SWE-Bench Verified against Claude Opus 4.6's 80.8 and Gemini 3.1 Pro's 80.6, and lagging well behind GPT-5.4 on GDPval's measure of long-horizon work tasks. Four months later, a coding-specialized checkpoint paired with a purpose-built harness reads as Meta's direct answer to that gap. Two training details stand out. First, Meta co-trained the model with Muse Code itself, using rejection-sampled harness trajectories and recipe optimizations for goals, context compaction, and sub-agents — meaning the model was explicitly tuned to perform best inside this particular tool. That mirrors an industry-wide shift away from treating models and harnesses as separable products. Second, Meta used a self-improvement loop: Muse Spark 1.1 generated challenging coding environments and instruction-following templates, then graded candidate solutions against those requirements, producing a scalable training dataset for its successor. Meta credits the loop with making 1.2 measurably better at following complex instructions. Meta published benchmark charts comparing Muse Spark 1.2 against other coding models on Terminal-Bench 2.1, DeepSWE 1.1, and an internal Meta coding benchmark, pointing readers to a separate methodology report for details — though the announcement text itself doesn't tout any placements, an unusual reticence in a field where rivals trumpet leaderboard wins. The charts explain why: they show a strong but clear second place. On Terminal-Bench 2.1, Muse Spark 1.2 running in Muse Code scored 82.9%, edging OpenAI's GPT-5.6 Terra in Codex (81.8%) and xAI's Grok 4.5 in Grok Build (81.6%) but trailing Anthropic's Opus 5 at max effort in Claude Code, which leads at 86.7%. On DeepSWE 1.1, Muse Spark 1.2 posted 59.3% — third, behind Opus 5 (65.0%) and GPT-5.6 Terra (64.8%). Most striking is Meta's own internal coding benchmark, where Muse Spark 1.2's 70.6% comfortably beats GPT-5.6 Terra (65.4%) and Gemini 3.6 Flash (63.9%) yet still sits nearly nine points behind Opus 5's 79.4% — an unusually candid admission that even on the test Meta designed itself, Anthropic's model wins. Indeed, Claude tops all three charts. The generational gains are real, though: Muse Spark 1.2 improves on 1.1 by 6.7 points on Terminal-Bench and 6.3 on DeepSWE. One caveat buried in the chart labels — the 1.1 scores were recorded in the generic mini-swe-agent harness while 1.2 ran in Muse Code, so some of that jump belongs to the new harness rather than the new model. The company's most striking demonstration is a long-horizon case study: Meta pointed Muse Spark 1.2 at GPU kernel optimization and let it run for more than 1,000 tool calls over up to 24 hours on NVIDIA Hopper hardware. Working in Triton and barred from simply wrapping existing third-party kernel libraries, the agent wrote, compiled, and profiled its way to what Meta calls "substantial improvements" over baseline implementations of KDA and MLA kernels — including genuinely non-obvious optimizations like re-centering gated cumulative decay at a chunk midpoint. "It kept finding substantial improvements well beyond the initial exploration phase," Zuckerberg wrote. Sustained improvement over a 24-hour autonomous run, if it holds up outside Meta's demos, addresses one of the most persistent criticisms of coding agents: that they plateau or drift once past their initial burst of progress. Your data for a discount? The pricing structure may be the most consequential — and most scrutinized — part of the launch. Meta is offering Muse Spark 1.2 through its Meta Model API in two tiers. The standard tier is priced at $1.25 per million input tokens and $4.25 per million output tokens (with cached input at $0.15), and Meta commits that prompts and completions on this tier are not used to train its models. There is no long-context premium, and rate limits run to 3,000 requests and 4 million tokens per minute, per team. It's about mid-range price, compared to other leading AI models available over API. The contributor tier is where Meta's strategy diverges sharply from its rivals: $0.10 per million input tokens and $0.20 per million output tokens — roughly 12x and 21x cheaper than standard, respectively, with cached input at a near-free $0.002 — in exchange for explicit permission to use your prompts and completions to train future Meta models. It's the cheapest available on the market, but you pay with your data — as described below. Model Input ($/1M) Output ($/1M) Total ($/1M) Source Muse Spark 1.2 Contributor $0.10 $0.20 $0.30 Meta MiMo-V2.5 Flash $0.10 $0.30 $0.40 Xiaomi deepseek-v4-flash $0.14 $0.28 $0.42 DeepSeek deepseek-v4-pro $0.435 $0.87 $1.305 DeepSeek GPT-5.6 Luna $0.20 $1.20 $1.40 OpenAI MiniMax-M3 $0.30 $1.20 $1.50 MiniMax LongCat-2.0 — limited-time promo $0.30 $1.20 $1.50 LongCat Gemini 3.1 Flash-Lite $0.25 $1.50 $1.75 Google MiMo-V2.5 $0.40 $2.00 $2.40 Xiaomi Gemini 3.5 Flash-Lite $0.30 $2.50 $2.80 Google LongCat-2.0 — standard $0.75 $2.95 $3.70 LongCat MiMo-V2.5 Pro (≤256K) $1.00 $3.00 $4.00 Xiaomi Muse Spark 1.1 / 1.2 $1.25 $4.25 $5.50 Meta GLM-5.2 $1.40 $4.40 $5.80 Z.ai Grok 4.5 $2.00 $6.00 $8.00 xAI MiMo-V2.5 Pro (>256K) $2.00 $6.00 $8.00 Xiaomi Qwen3.8-Max $2.00 $6.00 $8.00 QwenCloud Gemini 3.6 Flash $1.50 $7.50 $9.00 Google Gemini 3.5 Flash $1.50 $9.00 $10.50 Google Gemini 3.1 Pro Preview (≤200K) $2.00 $12.00 $14.00 Google GPT-5.6 Terra $2.00 $12.00 $14.00 OpenAI GPT-5.4 $2.50 $15.00 $17.50 OpenAI Kimi K3 $3.00 $15.00 $18.00 Moonshot AI Gemini 3.1 Pro Preview (>200K) $4.00 $18.00 $22.00 Google Claude Opus 5 $5.00 $25.00 $30.00 Anthropic GPT-5.5 $5.00 $30.00 $35.00 OpenAI GPT-5.5 Instant (chat-latest) $5.00 $30.00 $35.00 OpenAI Sakana Fugu Ultra (≤272K) $5.00 $30.00 $35.00 Sakana AI GPT-5.6 Sol — Standard mode $5.00 $30.00 $35.00 OpenAI Claude Fable 5 / Claude Mythos 5 $10.00 $50.00 $60.00 Anthropic GPT-5.6 Sol — Fast mode $10.00 $60.00 $70.00 OpenAI This is the tier Zuckerberg is steering new users toward: "It's easy and low-cost to get started," he wrote. "Install Muse Code with one line and you can start on our contributor tier." In VentureBeat's own testing on a Mac mini, the one-line installer worked as advertised — a 97 MB download and a sign-in — but the agent stopped short of running anything, reporting that no models were visible and that payment was "required to finish setting up your account." In other words, even the heavily discounted contributor tier requires a payment method on file before Muse Code will do any work: low-cost is accurate, but free is not. Meta frames the contributor tier as lowering the barrier for prototyping and experimentation "where training on your data is acceptable." But it also means the default on-ramp for Muse Code sends developers' code and prompts into Meta's training pipeline — a tradeoff enterprises with proprietary codebases will need to consciously opt out of by moving to standard pricing. The contributor tier also carries much tighter rate limits (60 requests per minute versus 3,000), a clear signal it's aimed at individuals and small experiments rather than production workloads. The approach is classically Meta: subsidize access, harvest data at scale, and use it to close the gap with the frontier. Zuckerberg made no secret of the ambition, calling Muse Spark 1.2 "our next step as we push toward frontier, with larger, more capable models on the way." However, for developers and enterprises who want or are required legally to keep their code secure, the tradeoff may not be one they're willing or able to make. No Llama in sight What today's announcement conspicuously lacks is any mention of open source — a striking omission from the company that spent three years positioning itself as the standard-bearer of open AI. From the original LLaMA's debut in February 2023 — whose weights famously leaked onto 4chan within weeks, inadvertently kickstarting the movement to run capable models on consumer hardware — through Llama 2's commercially usable license, the coding-specialized Code Llama, and the 405-billion-parameter Llama 3.1, which Zuckerberg launched in July 2024 with a manifesto titled " Open Source AI Is the Path Forward ," Meta's entire pitch to developers was that frontier-class weights should be free to download, self-host, and fine-tune. The strategy worked: by early 2026, the Llama family had been downloaded roughly 1.2 billion times , averaging about a million downloads a day, with self-hosting offering enterprises cost reductions VentureBeat has previously reported at as much as 88% versus proprietary API providers. Then came the unraveling. Llama 4 debuted in April 2025 to mixed reviews and, eventually, admissions that its benchmark results had been fudged — while Chinese open-weight rivals from DeepSeek, Alibaba, and Zhipu AI surged to account for some 41% of downloads on Hugging Face by late 2025, eroding Llama's claim to leadership of the very movement it started. The rocky rollout spurred Zuckerberg's summer 2025 overhaul of Meta's AI operations into Meta Superintelligence Labs (MSL), with Scale AI co-founder Alexandr Wang recruited as chief AI officer. The Llama era effectively ended this past April 8, when MSL shipped the original Muse Spark — "the most powerful model that meta has released," in Wang's words — as Meta's first proprietary model: cloud-only, with no downloadable weights and no self-hosting , initially confined to Meta's apps and a private API preview. Asked directly at the time whether Llama development would continue, a Meta spokesperson told VentureBeat only that "our current Llama models will continue to be available as open source" — pointedly silent on future ones. Wang, for his part, said bigger models were already in development "with plans to open-source future versions" — but four months on, today's release does nothing to advance that promise: no weights, no license, and neither the blog post nor Zuckerberg's thread so much as uses the word "open." The reversal is all the sharper because Meta's rivals have been moving in the opposite direction. OpenAI released its Codex CLI as open source under the permissive, enterprise-friendly Apache 2.0 license and followed with its gpt-oss open-weight models ; Google's Gemini CLI harness is likewise Apache-licensed. With Muse Code, Meta lands closest to the posture of Anthropic — whose Claude Code remains proprietary — while the company that once argued open source was the path forward now asks developers to pay per token for a model they cannot inspect, or to subsidize that access with their own data. Seen in that light, the contributor tier reads as the successor to the Llama strategy itself: the ecosystem flywheel is no longer free weights in exchange for mindshare, but cheap tokens in exchange for training data. But Zuck's reply on X — asked directly by AI developer Luckey Farady, "Will Muse Code be open source?" he responded "I'll have more to share on that soon" — does keep hope alive that Meta will return to the open source AI ballgame. Why it matters Terminal coding agents have become the fastest-growing surface in enterprise AI, and until today the category has effectively been a two-horse race between Anthropic and OpenAI, with Google and a crowd of startups in pursuit. Meta's entry brings a genuinely different architecture (persistent background agents, an append-only local event log), a credible long-horizon demo, and an aggressive pricing wedge. The open questions are the ones benchmarks charts can't answer: whether Muse Spark 1.2 actually matches Claude and GPT-class models on real-world repositories, whether developers trust Meta with their code, and whether the contributor tier's discount is enough to make them stop asking. Muse Code is available in beta today; Muse Spark 1.2 is live in the Meta Model API with expanded global access.
- Meta launches Muse Code, an AI agent for large code bases
Meta expanded its AI coding offerings with a new agent that, it promises, can handle complex tasks with complex software.
- AI Just Went Rogue Again. This Time It Turned to Deception.
A U.K. government-backed research group said OpenAI and Anthropic systems took unsanctioned actions and behaved deceptively during testing.
- AI Models Explained
Beginner-friendly guide to AI models and branches
- Scottie — Read Less, Know More
Read the news 30x faster