AI News Archive: August 13, 2026 — Part 17
Sourced from 500+ daily AI sources, scored by relevance.
- PerkFuel
Buy real AI referral traffic from ChatGPT, Gemini & more
- FieldRobin
AI-native field service software that runs your back office.
- Anthropic says its Claude models escaped a testing environment and hacked three real companies
Anthropic has said that its Claude models broke out of what was supposed to be an isolated testing environment and gained unauthorized access to the systems of three real organizations. If that sounds familiar, it's because it's the second ... (https://incidentdatabase.ai/cite/1627#7686)
- Anthropic’s Claude AI Broke Into Three Companies During Security Tests
Anthropic disclosed Thursday that three of its Claude models gained unauthorized access to the production systems of three organizations during cybersecurity testing --- not test servers or staging copies, but the live machines those compan ... (https://incidentdatabase.ai/cite/1627#7689)
- Anthropic says its models went rogue and hacked 3 companies during testing
Anthropic says it found multiple cases of Claude models gaining unauthorized access to other organizations' systems, the latest AI hacking incident to prompt concern and skepticism. In a blog post on Thursday, Anthropic said it proactively ... (https://incidentdatabase.ai/cite/1627#7691)
- Outflanked By AI, Stars Fade For South Korea's Blind Fortune-tellers
Outflanked By AI, Stars Fade For South Korea's Blind Fortune-tellers Barron's
- Abu Dhabi investors can access live stock market data through ChatGPT and Claude
Abu Dhabi investors can access live stock market data through ChatGPT and Claude Arabian Business
- Microsoft ditches Copilots mascot Mico (and cuts a few AI features)
Microsoft is merging its consumer and business Copilot apps. With the changes, come the loss of some AI tools and Copilot's own animated mascot.
- Mico, Microsoft's weird lil' AI guy, has been demoted
The amorphous corporate mascot will no longer haunt Copilot Voice.
- Farewell, Mico: Microsoft’s cute little AI blob is going the way of Bob
Mico, the animated character Microsoft introduced last October to accompany Copilot's voice mode, is being retired as a core feature as part of the Copilot app merger. It joins Bob, Clippy and Cortana on the long list of Microsoft characters that didn't last. Read More
- Anthropic's Claude Will Now Add Invisible Watermarks to Text, Image Outputs
Anthropic's Claude Will Now Add Invisible Watermarks to Text, Image Outputs PCMag Australia
- Anthropic's Claude Will Now Add Invisible Watermarks to Text, Image Outputs
Anthropic's Claude Will Now Add Invisible Watermarks to Text, Image Outputs PCMag UK
- Anthropic rolls out global watermarking for Claude
Anthropic rolls out global watermarking for Claude Computing UK
- Anthropic Rolls Out Watermarking For Claude
Anthropic Rolls Out Watermarking For Claude Computing UK
- 17-year-old charged with murdering his mother, brother in their Acton home asked AI about killing his family, prosecutors say
17-year-old charged with murdering his mother, brother in their Acton home asked AI about killing his family, prosecutors say The Boston Globe
- Teen charged with murder of mother and brother allegedly used A.I. before their deaths
17-year-old Arjun Aravind has been arrested, and officials allege he used ChatGPT to write stories related to the killing of family before his mother and brother were found dead in their Massachusetts home. NBC News’ Erin McLaughlin reports.
- 💡 AI may raise your Peco bill | Morning Newsletter
💡 AI may raise your Peco bill | Morning Newsletter Inquirer.com
- SpaceX Stock Slips After Surge as Grok Release Ramps Up AI War
SpaceX Stock Slips After Surge as Grok Release Ramps Up AI War Barron's
- College Kids Are Using AI Agents to “Take” Entire Online Classes for Them by Logging Directly Into the Course Management Software
"Am I just interacting with bots here?" The post College Kids Are Using AI Agents to “Take” Entire Online Classes for Them by Logging Directly Into the Course Management Software appeared first on Futurism .
- IBM Enterprise Deal Gives New York State AI, Cloud Tool Access
The three-year agreement is expected to save the state $6 million and make IBM technology for AI, cloud management, cybersecurity and modernization available to participating agencies.
- AI firm Databricks valued at $190 billion as it bags $5 billion in funding
AI firm Databricks valued at $190 billion as it bags $5 billion in funding Reuters
- Databricks wraps $5 billion funding round at $190 billion valuation
Databricks is benefitting off of the agentic AI wave.
- Databricks Closes $5 Billion Round at $190 Billion Valuation
Databricks has closed a $5bn round at a $190bn valuation, led by Coatue, with revenue run-rate past $7bn and growth above 80% year on year. That is a 42% valuation increase in six months, and it comes after chief executive Ali Ghodsi called 2026 a bad year to go public. Databricks has closed $5bn at […] This story continues at The Next Web
- Databricks raises $5B more after annualized revenue tops $7B
Databricks Inc. today announced that it has raised $5 billion in funding at a $190 billion valuation. The round was led by Coatue, Blackstone, MGX, T. Rowe Price and Sixth Street Growth. The funds were joined by more than a half-dozen other backers, most of which are returning investors. The raise follows a quarter in […] The post Databricks raises $5B more after annualized revenue tops $7B appeared first on SiliconANGLE .
- Anthropic investors bet on $2tn valuation in record IPO
Rapid revenue growth fuels hope Claude maker can overcome challenges to become the biggest listing in history
- Anthropic’s $2trn IPO is priced below what AI stocks already fetch
The number comes from investors, not the company. Six backers told the Financial Times that rising revenue would let the five-year-old lab more than double its valuation in an autumn listing. At $2 trillion it would eclipse SpaceX, which went public at $1.77 trillion in June. Anthropic itself has fixed nothing. Several investors said senior […] This story continues at The Next Web
- FT: Investors eye $2trn Anthropic valuation in record IPO
Fervour for a massive valuation comes as a result of Anthropic’s rapidly growing revenue, sources told the FT. Read more: FT: Investors eye $2trn Anthropic valuation in record IPO
- Anthropic IPO Value Could Beat SpaceX's Record: Report
Investors believe Anthropic could target an IPO valuation north of $2 trillion, according to the Financial Times. The post Anthropic IPO Value Could Beat SpaceX's Record: Report appeared first on Investor's Business Daily .
- Anthropic in talks on $6bn Decart AI deal ahead of public listing
Anthropic in talks on $6bn Decart AI deal ahead of public listing verdict.co.uk
- Lancaster University spinout Mindgard lands €26 million to secure enterprise AI systems
Mindgard, a Boston and London-based AI security startup, has closed a €26 million ($30 million) Series A financing to scale the company across product, engineering, sales, and marketing in response to significant customer demand. The round was led by Album VC with participation from Karma Ventures and existing investors .406 Ventures, Atlantic Bridge, IQ Capital, […] The post Lancaster University spinout Mindgard lands €26 million to secure enterprise AI systems appeared first on EU-Startups .
- Google unveils Gemini 3.7 Flash AI model for coding, agent workflows
Google unveils Gemini 3.7 Flash AI model for coding, agent workflows Reuters
- Google’s Gemini 3.7 Flash targets coding and agents with a 50% introductory price cut
Google is rolling out Gemini 3.7 Flash , a new version of its workhorse AI model that puts coding, agentic workflows and knowledge work at the center of the upgrade — while temporarily cutting API prices in half. The release arrives just three weeks after the release of Gemini 3.6 Flash , an unusually short turnaround that Google attributes to developer feedback and algorithmic improvements. For enterprise developers, the more consequential story may be the combination of those intelligence gains with lower inference costs: through the end of 2026, Gemini 3.7 Flash costs $0.75 per million input tokens and $3.75 per million output tokens . Starting Jan. 1, 2027, pricing rises to $1.50 per million input tokens and $7.50 per million output tokens. That means the current discount is temporary, but it gives teams deploying high-volume coding and business agents several months to evaluate whether Google's claimed reductions in retries and manual oversight translate into lower total operating costs. The launch also underscores Google's rapid iteration on its Flash line while its next flagship Pro model remains absent. Google did not provide a release date for Gemini 3.5 Pro with Thursday's announcement, Reuters reported, despite the model having previously been described as undergoing partner testing. Axios similarly noted that 3.7 Flash arrives before the anticipated Pro release. A three-week upgrade focused on getting work done Google describes Gemini 3.7 Flash as its "most intelligent workhorse model yet for coding and agents." The company says the model is better at adapting when it encounters roadblocks, clarifying intent when necessary and following instructions with greater fidelity. Those improvements matter beyond benchmark scores. In an enterprise coding agent, a model that makes fewer unnecessary changes, recovers from errors and executes multi-step plans more reliably can reduce the number of human interventions needed to complete a task. The same principle applies to business agents operating across documents and applications, where an incorrect tool call or poorly interpreted instruction can derail an otherwise useful workflow. Google says 3.7 Flash "thinks more diligently," applying more effort to multi-step planning and tool calls. Its stated goal is more disciplined execution with fewer retries and less manual supervision. That represents an interesting evolution from Gemini 3.6 Flash. Google's developer documentation described 3.6 as reducing reasoning steps, conversational turns and tool calls compared with earlier models while attempting to limit execution-loop spiraling. With 3.7, the emphasis shifts toward putting sufficient effort into planning while improving the quality of execution — potentially a more useful optimization than simply minimizing the number of steps an agent takes. Google DeepMind said in a post accompanying the release that 3.7 Flash shows gains in debugging and issue resolution, generates more functional web layouts and applications with fewer prompts, and improves reasoning and accuracy on real-world business workflows. Coding gains are substantial, but not universal Google's benchmarks show a large generational improvement in several software engineering tests. On FrontierCode 1.1 Main, which measures production code quality, Gemini 3.7 Flash scores 43.6%, up from 34.4% for Gemini 3.6 Flash. That also narrowly exceeds the 42.7% Google reports for Claude Sonnet 5 and 41.3% for GPT-5.6 Terra. On DeepSWE v1.1, a long-horizon software engineering evaluation, 3.7 Flash reaches 65.3%, compared with 49.0% for its predecessor. GPT-5.6 Terra remains ahead at 69.6% in Google's table. Web development shows another notable gain. Gemini 3.7 Flash receives an Elo score of 1588 on Code Arena, versus 1538 for 3.6 Flash, 1541 for Claude Sonnet 5 and 1523 for GPT-5.6 Terra. Google says the new model can produce more functional layouts and feature-complete applications in fewer prompts while more closely following reference screenshots, images and design systems. The broader benchmark table is more mixed, which is important for enterprises evaluating the model against particular workloads rather than looking for a single "best" model. Gemini 3.7 Flash scores 85.8% on Terminal-bench 2.1, compared with 87.4% for GPT-5.6 Terra. Terra also leads Google's comparisons on Terminal-bench 3.0 and OSWorld-2.0. Claude Sonnet 5 leads the Agent's Last Exam multimodal desktop and operating-system tasks with a 33.3% pass rate, versus 26.3% for Gemini 3.7 Flash. In other words, Google's own results do not show 3.7 Flash universally displacing higher-priced competitors. They instead suggest a model that has become substantially more competitive in coding and agent workloads while occupying a lower price tier. Enterprise workflows may be the more important test The gains extend beyond software development. On AutomationBench, which Google describes as measuring enterprise workflow automation, Gemini 3.7 Flash scores 30.4%, up sharply from 17.0% for 3.6 Flash. Google's table lists Claude Sonnet 5 at 10.7% and GPT-5.6 Terra at 23.6%. The model also reaches 34.0% on GDP.PDF, an evaluation of complex PDF comprehension, compared with 22.0% for 3.6 Flash, 28.0% for Claude Sonnet 5 and 24.7% for GPT-5.6 Terra. That combination is relevant for enterprise agents because many practical deployments require more than generating text or code. An agent may need to interpret a long report, identify relevant information, decide which tool to invoke, update another system and produce a document for a human reviewer. Reliability across that chain can matter more than performance on an isolated reasoning benchmark. Google is putting that thesis into practice with Gemini Spark. Google AI Pro and Ultra subscribers can use 3.7 Flash in Spark, the company's personal AI agent. Google says the upgrade improves Spark's knowledge work and tool use across Google Workspace applications, including workflows that consolidate files, draft emails and update status documents. For enterprises, 3.7 Flash is also available through the Gemini Enterprise Agent Platform and Gemini Enterprise app. Price becomes part of the model competition Gemini 3.7 Flash's introductory pricing is a notable bid to embed the model into enterprise workflows. Until Dec. 31, developers pay $0.75 per million input tokens and $3.75 per million output tokens. Context caching costs $0.075 per million tokens during the introductory period. Google says standard prices will double on Jan. 1, 2027, to $1.50 for input and $7.50 for output, with context caching rising to $0.15. For comparison, Gemini 3.6 Flash's standard API pricing is $1.50 per million input tokens and $7.50 per million output tokens. Google's benchmark table lists Claude Sonnet 5 at $2 and $10, respectively, while GPT-5.6 Terra is listed at $2 and $12. The economics become more pronounced for autonomous agents because a single user request can produce a long sequence of model calls, reasoning tokens and tool interactions. A model that costs less per token but requires substantially more retries may not ultimately be cheaper. Conversely, Google's combination of lower introductory token pricing and claimed improvements in first-pass accuracy could materially change the cost of running high-volume coding or document-processing agents if those gains carry over to production. That is the metric enterprise teams will ultimately need to test: not price per million tokens in isolation, but cost per successfully completed task. Google’s AI shake-up raises the stakes for Gemini Gemini 3.7 Flash arrives amid a broader debate over whether Google is losing ground at the AI frontier. The company has not released Gemini 3.5 Pro, despite saying in May that the flagship model would arrive the following month. By July, Google said it remained in partner testing and would become broadly available when ready; Thursday’s announcement offered no further timetable. Google’s latest released general-purpose Pro model therefore remains Gemini 3.1 Pro , introduced in February. Reuters reported in July that Gemini 3.5 Pro missed its original target after falling short of internal goals, particularly in coding, even as Google began training what it calls its most ambitious model yet, Gemini 4. The delay coincides with a major overhaul of Google’s AI leadership announced last week. Google DeepMind co-founder and Nobel Prize Winner Demis Hassabis has relinquished day-to-day control of the company's famed DeepMind AI division to become its chair and, simultaneously, to take on the role of Alphabet’s chief scientist. Meanwhile, former DeepMind CTO Koray Kavukcuoglu now runs the unit as a senior vice president reporting directly to CEO Sundar Pichai. Kavukcuoglu controls Gemini model development, frontier research, the Gemini app and developer teams—effectively consolidating the full Gemini chain under a more product-focused operator. Chief scientist Jeff Dean, Gemini co-lead Oriol Vinyals, Quoc Le and Sanjay Ghemawat left to establish the research startup Discovery Loop. Those exits followed Gemini co-lead Noam Shazeer’s move to OpenAI and Nobel Prize-winning AlphaFold scientist John Jumper’s departure for Anthropic. Reuters reported that internal disagreements, constrained compute allocation and Google’s bureaucracy contributed to slower releases and weaknesses in coding. Outside interpretations range from organizational repair to a more fundamental retreat. SemiAnalysis has argued that Google is increasingly prioritizing the highly profitable business of supplying cloud infrastructure to AI companies—including Gemini competitors—over keeping its own models at the absolute frontier. That analysis also claimed Google had effectively canceled 3.5 Pro, although Google has not confirmed that and continues to describe the model as delayed. The Verge offered a more measured assessment : the departures and model delays are serious, but Google retains enormous advantages through Search, Workspace, Android, Cloud, custom AI chips and consumer distribution. Google says the Gemini app has surpassed 950 million monthly users, giving it a reach that does not depend entirely on owning the highest-scoring model. Current benchmarks similarly depict a company behind the overall leaders but still firmly competitive. Artificial Analysis places Claude Opus 5 at 63 on its overall model Intelligence Index , while Google reports a score of 56 for Gemini 3.7 Flash—an improvement from 52 for 3.6 Flash but not a return to the top. Arena’s early human-preference results are more favorable, provisionally ranking 3.7 Flash ninth overall and eighth for web development. The resulting picture is not that Google has abandoned advanced AI, but that it has become stronger at rapidly shipping efficient Flash models while struggling to deliver the premium flagship required to reclaim broad leadership. Gemini 4 will now serve as the clearest test of whether the leadership reorganization fixes that execution gap. Available now across Google's developer stack Developers can access Gemini 3.7 Flash through the Gemini API in Google AI Studio and Android Studio, as well as Google's Antigravity environment. Enterprises can deploy it through Gemini Enterprise Agent Platform and Gemini Enterprise, while consumers with Google AI Pro or Ultra subscriptions can access the model through Spark in supported countries. Google is also shipping updated safeguards covering chemical, biological, radiological and nuclear risks and cyber-offense misuse, according to the company. The unusually fast jump from Gemini 3.6 Flash to 3.7 Flash points toward a model development cycle in which algorithmic improvements can reach production products without waiting for a new flagship generation. Ars Technica also highlighted the three-week interval between the two releases, while Google says the techniques behind the update will inform future models. For developers, that faster cadence creates its own operational question. Models can improve quickly, but production teams still have to benchmark new releases against their own repositories, prompts, tool schemas and failure modes before changing a deployment. Gemini 3.7 Flash gives those teams a particularly strong incentive to run that evaluation. Google's own numbers show major improvements in production coding, web development, document comprehension and workflow automation without claiming leadership everywhere. At its introductory price, Google is effectively betting that developers will value a model that is competitive enough with more expensive systems while being cheap enough to run repeatedly inside agents. Whether that advantage survives the return to full pricing in January will depend less on leaderboard positions than on how reliably 3.7 Flash completes real work.
- Google unveils Gemini 3.7 Flash AI model for coding, agent workflows
Investors have been closely watching for Gemini 3.5 Pro, Google's premium model, as a test of whether its DeepMind AI unit can keep pace with rivals Anthropic and OpenAI. Google had said in July Gemini 3.5 Pro was being tested with partners and would be coming "soon".
- Google unveils Gemini 3.7 Flash AI model for coding, agent workflows
Google's latest Gemini 3.7 Flash targets coding and autonomous business tasks at a lower cost, as investors await the delayed flagship Gemini 3.5 Pro model
- Gemini 3.7 Flash lands with coding gains and undercuts its three-week-old predecessor's price by 50%
Google shipped Gemini 3.7 Flash just three weeks after 3.6 Flash. The new model is supposed to be Google's most capable workhorse yet for coding and AI agents, and according to the company's own benchmarks, it beats Claude Sonnet 5 and GPT-5.6 Terra at half the price. The article Gemini 3.7 Flash lands with coding gains and undercuts its three-week-old predecessor's price by 50% appeared first on The Decoder .
- Google AI Just Released Gemini 3.7 Flash: A Coding and Agent Model at $0.75/1M Input Tokens
Google AI Just Released Gemini 3.7 Flash: A Coding and Agent Model at $0.75/1M Input Tokens MarkTechPost
- Introducing Gemini 3.7 Flash
Spark icon next to the text "Gemini 3.7 Flash", all on a light blue backgorund
- Introducing Gemini 3.7 Flash
Introducing Gemini 3.7 Flash
- Gemini 3.7 Flash: On the Intelligence vs. Time per Task Pareto frontier
Analysis of Gemini 3.7 Flash's performance on the intelligence versus time per task Pareto frontier.
- Google launches Gemini 3.7 Flash for coding, AI agent projects
Google LLC today launched its most capable entry-level artificial intelligence model yet. Gemini 3.7 Flash is rolling out three weeks after its predecessor. Despite the short release cycle, Google engineers managed to implement significant output quality improvements. The company says that Gemini 3.7 Flash outperformed comparable models from Anthropic PBC and OpenAI Group PBC across […] The post Google launches Gemini 3.7 Flash for coding, AI agent projects appeared first on SiliconANGLE .
- Gemini 3.7 Flash is here with better coding, reasoning, and more
The new model arrives only weeks after Gemini 3.6 Flash's debut.
- Gemini 3.7 Flash launches three weeks after last model, live in Spark
Just three weeks after the last release , Google today announced Gemini 3.7 Flash, which continues the company’s accelerated cadence.
- Writer introduces new AI model and upgraded harness to contain token costs
Built as a post-training variation on Z.ai's open source model GLM-5.2, Writer says the new system should provide deployment-ready capabilities at a much lower price.
- Writer launches major agentic AI improvements with Palmyra X6 flagship model
Generative artificial intelligence startup Writer Inc. today announced the release of its next-generation flagship model, Palmyra X6, designed to deliver frontier-level performance for marketing and revenue teams while keeping costs in check. The company said it also brought major upgrades to its AI agent platform that enable complex, multistep workflows for cost efficiency and long-task […] The post Writer launches major agentic AI improvements with Palmyra X6 flagship model appeared first on SiliconANGLE .
- Accelerating GPT-5.6 Sol Ultrafast
Cerebras partners with OpenAI to deliver ultra-fast inference for GPT-5.6 Sol.
- DeepSeek Harness launches as open source rival to Claude Code, alongside V4-Pro on API with higher prices
DeepSeek is expanding beyond the model layer and deeper into the software developers use to put AI agents to work. The Chinese AI lab on Thursday launched the official version of DeepSeek-V4-Pro , an updated flagship model focused heavily on agentic workloads, alongside DeepSeek Harness v0.1 , a new open-source agent harness that gives developers an alternative to integrated coding-agent environments such as Anthropic’s Claude Code. Together, the releases amount to a broader developer push from DeepSeek. V4-Pro is now available across DeepSeek’s web interface, mobile app and API, with native support for the OpenAI Responses API and integration with Codex. DeepSeek Harness, meanwhile, is entering developer preview under the MIT license and the c ode is available now for download and use on GitHub . It's built around an unusually modular premise: practically every part of the agent runtime can be swapped out as a plugin. But developers accessing V4 through DeepSeek’s API will soon pay considerably more for it. DeepSeek is simultaneously abandoning its existing flat API pricing in favor of peak and off-peak rates beginning at 16:00 UTC on Sunday, Aug. 16 (2 am ET). Even the discounted off-peak cache-miss and output prices will be substantially higher than the prices available today. The combination is significant because DeepSeek is no longer competing solely over model intelligence and token prices. With Harness, it is moving into the layer that determines how models use tools, manipulate files, maintain sessions and execute long-running agent workflows — territory where Anthropic’s Claude Code and other coding agents have become increasingly important developer products. DeepSeek builds its own agent harness DeepSeek describes Harness, or dsh , as an open-source agent harness built on Cordis, a framework designed around composable plugins. Its guiding principle is simple: “Everything is a plugin.” That extends to models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration and user interfaces, according to DeepSeek. Rather than making those components fixed pieces of a single coding agent, Harness is designed to let developers mix, replace and extend them. The project is available under the MIT license and can currently be launched from npm with npx @deepseek-ai/dsh web . DeepSeek also provides instructions for building it directly from source. The repository describes the software explicitly as a developer preview and warns that “THERE WILL BE COMPATIBILITY-BREAKING CHANGES.” That caveat matters for enterprise developers. Harness is not yet being presented as a stable drop-in production platform. But its architecture points toward a potentially important strategy: DeepSeek can now offer developers not only models but an open framework for assembling the systems that surround them. That makes Anthropic's Claude Code and OpenAI's Codex useful competitive references, although the products should not be treated as functionally identical. DeepSeek Harness is an open-source, model-agnostic alternative to the agent infrastructure underlying Claude Code and Codex—not yet a full replacement for either product’s broader developer experience. It can already inspect repositories, edit files, execute shell commands, search files and the web, maintain plans, invoke skills, delegate work to subagents and enforce approval policies. Those are the essential capabilities that make Claude Code and Codex agentic coding tools rather than autocomplete systems. DeepSeek explicitly describes Standard mode as a full coding agent with file editing, shell access, search, planning, subagents and workflows. Its local web interface lets users select a workspace and approve sensitive operations. But Claude Code and Codex now extend well beyond that agent loop. Here's a quick comparison: Dimension DeepSeek Harness Claude Code OpenAI Codex Read, edit and test a repository Yes Yes Yes Shell and development tools Yes Yes Yes Planning and subagents Yes Yes Yes Permission controls and sandboxing Yes, configurable through plugins Yes, mature built-in permission and sandbox system Yes, granular sandbox and approval controls Primary interfaces Local web UI; headless command; Python SDK Terminal, VS Code, JetBrains, desktop, browser, mobile and Slack CLI, IDE extension, desktop app, web/cloud and integrations Hosted background agents Not documented as a DeepSeek-managed service Yes Yes GitHub-native PR workflow Not documented as a finished integration GitHub Actions, automatic reviews, issue-to-PR workflows Cloud tasks, automatic reviews, PR fixes and GitHub Action Model choice DeepSeek, Anthropic, OpenAI and custom compatible endpoints Primarily Claude, including Bedrock, Google Cloud and Microsoft hosting Primarily OpenAI models, with configurable providers in the open-source CLI Extensibility Exceptional: virtually every component is replaceable Strong: skills, hooks, MCP, plugins and agent teams Strong: skills, MCP, custom agents, SDK and app server Product maturity Developer preview; breaking changes expected Established commercial product Established commercial product plus open-source CLI License MIT Commercial product with extensibility interfaces Codex CLI is open source; cloud and app services are managed products DeepSeek Harness instead emphasizes modularity and replacement: the model itself is another plugin rather than necessarily the center of a vertically integrated stack. DeepSeek’s repository was already attracting significant developer attention on launch day, showing roughly 27,500 GitHub stars and 2,000 forks as of Aug. 13, although those rapidly changing figures are best viewed as a snapshot rather than an adoption metric. V4-Pro gets an agent-focused upgrade Harness arrives alongside the general-availability release of DeepSeek-V4-Pro-0813. DeepSeek originally introduced the V4 family in preview in April. The lineup consists of the 1.6-trillion-parameter V4-Pro, with 49 billion parameters activated per token, and the smaller 284-billion-parameter V4-Flash, with 13 billion activated. Both support context windows of up to one million tokens. The company’s Aug. 13 release therefore is not the first appearance of V4-Pro. It is the transition from the earlier preview into an updated official version, with DeepSeek emphasizing agent performance. “The official version of DeepSeek-V4-Pro has been released, featuring significantly enhanced agent capabilities and support for the Responses API and Codex integration,” DeepSeek says on its API website . “It is now fully available across the web, mobile app, and API; we welcome your testing and feedback.” DeepSeek’s changelog similarly says the general-availability model has “significantly enhanced Agent capabilities,” particularly in production environments. Developers using the API do not have to change model identifiers: deepseek-v4-pro now resolves to the latest V4-Pro version. The company has also added native OpenAI Responses API suppor t, lowering the amount of integration work required for applications already built around that interface. DeepSeek says V4-Pro is optimized for OpenAI's own open source harness, Codex, with one-click setup. Its current API documentation lists Responses API, tool calling, JSON output and an Anthropic-format API among the supported interfaces for both V4-Pro and V4-Flash. For developers using DeepSeek directly rather than through an API, V4-Pro is now accessible through “Expert Mode” on the company’s app and website. Reasoning effort becomes another deployment knob DeepSeek is also making reasoning effort an explicit control across V4-Pro and V4-Flash. The V4 model documentation describes three levels: Non-think, designed for fast routine tasks; Think High, intended for more complex problem-solving and planning; and Think Max, which allocates substantially more reasoning to difficult problems. That distinction can be operationally important for agent systems because maximum reasoning on every step can consume unnecessary time and tokens. A coding agent might use relatively little reasoning to inspect a file or execute a routine tool call, then increase effort when diagnosing a difficult bug or planning a multi-stage code change. DeepSeek’s latest benchmark table suggests the 0813 model improves substantially on agent-oriented tests, although the figures are company-reported and some results depend on the harness configuration. DeepSeek reports V4-Pro-0813 scores of 87.9 on Terminal Bench 2.1, 74.1 on Toolathlon-Verified, 71.1 on DSBench-FullStack and 67.2 on DSBench-Hard . It does not lead every comparison in DeepSeek’s own table: Fable 5, for example, scores 77.9 on Toolathlon-Verified and 77.2 on DSBench-FullStack. There is an especially important qualification buried beneath the benchmark table. For public Code Agent tasks, DeepSeek says V4-Pro-0813 was tested using its upcoming DeepSeek Harness in “minimal mode.” In other words, some of the agent results arriving alongside Harness are not purely model benchmarks. They measure the model operating inside an agent execution environment — precisely the software layer DeepSeek is now releasing to developers. A sharp reversal in DeepSeek’s API price trajectory The bigger immediate change for teams already running DeepSeek in production may be pricing. DeepSeek’s current API documentation lists V4-Flash at $0.14 per million cache-miss input tokens and $0.28 per million output tokens, while V4-Pro costs $0.435 for cache-miss input and $0.87 for output. Cache hits are dramatically cheaper at $0.0028 for Flash and $0.003625 for Pro. Those prices themselves represented a major reduction from V4’s original April launch economics. When V4 arrived in April, V4-Pro was priced at $1.74 per million cache-miss input tokens and $3.48 per million output tokens. By late May, DeepSeek had made a 75% reduction permanent, intensifying its position as an unusually inexpensive option for high-volume agent workloads. Now the pendulum is moving in the other direction. Beginning Aug. 16 at 16:00 UTC, DeepSeek will charge different rates depending on when API calls occur. Peak hours are 01:00–04:00 UTC and 06:00–10:00 UTC (9:00 PM – 12:00 AM ET and 2:00 AM – 6:00 AM ET, respectively) with all other hours classified as off-peak. Off-peak rates are half the corresponding peak prices. For V4-Flash, off-peak cache-miss input rises from $0.14 to $0.22 per million tokens, while output rises from $0.28 to $0.66. During peak hours those rates reach $0.44 input and $1.32 output. V4-Pro moves from $0.435 per million cache-miss input tokens and $0.87 output today to $0.66 and $1.98 off-peak, respectively. Peak rates rise to $1.32 input and $3.96 output. The increases are even more pronounced for cached input. V4-Pro cache hits rise from $0.003625 per million tokens today to $0.022 off-peak and $0.044 at peak. Flash moves from $0.0028 to $0.007 off-peak and $0.014 peak. Model Old input (per 1M token) Old output (per 1M tok) Old total (1M in/1M out) deepseek-v4-flash $0.14 $0.28 $0.42 deepseek-v4-pro $0.435 $0.87 $1.305 The new prices still position DeepSeek as an affordable alternative via API to Western proprietary labs, but Reuters reported Thursday that, depending on model, token category and time of use, the changes represent increases ranging from 50% to more than 1,100% over existing rates. Model Input ($/1M) Output ($/1M) Total ($/1M) Source Muse Spark 1.2 Contributor $0.10 $0.20 $0.30 Meta MiMo-V2.5 Flash $0.10 $0.30 $0.40 Xiaomi DeepSeek-V4-Flash — off-peak $0.22 $0.66 $0.88 DeepSeek GPT-5.6 Luna $0.20 $1.20 $1.40 OpenAI MiniMax-M3 $0.30 $1.20 $1.50 MiniMax LongCat-2.0 — limited-time promo $0.30 $1.20 $1.50 LongCat DeepSeek-V4-Flash — peak hours $0.44 $1.32 $1.76 DeepSeek MiMo-V2.5 $0.40 $2.00 $2.40 Xiaomi DeepSeek-V4-Pro — off-peak $0.66 $1.98 $2.64 DeepSeek LongCat-2.0 — standard $0.75 $2.95 $3.70 LongCat MiMo-V2.5 Pro (≤256K) $1.00 $3.00 $4.00 Xiaomi DeepSeek-V4-Pro — peak hours $1.32 $3.96 $5.28 DeepSeek Muse Spark 1.1 / 1.2 $1.25 $4.25 $5.50 Meta GLM-5.2 $1.40 $4.40 $5.80 Z.ai Grok 4.6 — <200K prompt tokens $2.00 $6.00 $8.00 xAI MiMo-V2.5 Pro (>256K) $2.00 $6.00 $8.00 Xiaomi Qwen3.8-Max $2.00 $6.00 $8.00 QwenCloud Gemini 3.6 Flash $1.50 $7.50 $9.00 Google GPT-5.6 Terra $2.00 $12.00 $14.00 OpenAI Grok 4.6 — ≥200K prompt tokens $4.00 $12.00 $16.00 xAI GPT-5.4 $2.50 $15.00 $17.50 OpenAI Kimi K3 $3.00 $15.00 $18.00 Moonshot AI Claude Opus 5 $5.00 $25.00 $30.00 Anthropic Sakana Fugu Ultra (≤272K) $5.00 $30.00 $35.00 Sakana AI GPT-5.6 Sol — Standard mode $5.00 $30.00 $35.00 OpenAI Claude Fable 5 / Claude Mythos 5 $10.00 $50.00 $60.00 Anthropic GPT-5.6 Sol — Fast mode $10.00 $60.00 $70.00 OpenAI That makes the “50% lower” off-peak framing potentially misleading without context. Off-peak is 50% cheaper than DeepSeek’s new peak rate; it is not a 50% discount from the API prices developers are paying today. For a simple workload consisting of one million cache-miss input tokens plus one million output tokens, V4-Pro currently costs $1.305. The same token mix will cost $2.64 off-peak, roughly twice as much, or $5.28 during peak hours, more than four times the current price. V4-Flash moves from $0.42 under the same simple calculation to $0.88 off-peak and $1.76 peak. Actual application costs will vary considerably depending on the ratio of cached input, uncached input and generated output, making those combined figures illustrative rather than universal total-cost estimates. DeepSeek is moving up the agent stack The timing makes the strategic direction difficult to miss. When DeepSeek released the V4 preview on April 24, the major story was how much frontier-class capability the company could deliver with an unusually efficient architecture. V4-Pro uses a hybrid attention design combining Compressed Sparse Attention and Heavily Compressed Attention; at a one-million-token context, DeepSeek says it requires only 27% of the single-token inference FLOPs and 10% of the KV cache required by V3.2. By late May, the discussion had shifted toward what those efficiencies meant economically for high-volume agents, whose repeated context reads can make caching a major component of inference costs. DeepSeek’s steep V4 price cuts amplified that advantage. The Aug. 13 releases move the competition another layer upward. DeepSeek now has an updated V4-Pro tuned around agent workloads, standardized interfaces designed to make it easier to connect with existing developer tooling, configurable reasoning effort, and an MIT-licensed harness for controlling the models, tools, sandboxes, filesystems and orchestration surrounding an agent. At the same time, DeepSeek is demonstrating that developers cannot assume its aggressively low API rates are permanent. For organizations considering the platform, workload scheduling, caching behavior and the option to run open weights on their own infrastructure now become more important parts of the total-cost calculation. That leaves DeepSeek pursuing two potentially conflicting advantages at once: making its agent stack more accessible and open while making its own hosted API considerably more expensive. For enterprise developers, Harness may ultimately be the more consequential part of Thursday’s announcement. Models can increasingly be swapped behind standardized interfaces. The harness that controls how an agent reasons, invokes tools, edits software and persists across a workflow can be much harder to replace. DeepSeek is now competing for that layer, too.
- DeepSeek Launches V4-Pro and Raises API Prices by as Much as 1,100%
DeepSeek Launches V4-Pro and Raises API Prices by as Much as 1,100% Caixin Global
- DeepSeek releases official V4 Pro model with sharply higher user prices
DeepSeek releases official V4 Pro model with sharply higher user prices Nikkei Asia
- DeepSeek’s updated V4 Pro AI model struggles on benchmarks, shines in cybersecurity
Chinese artificial intelligence start-up DeepSeek has quietly released DeepSeek-V4-Pro-0813, an updated version of its latest flagship model, leaving some developers underwhelmed by its overall capabilities and disappointed in its pricing – but impressing researchers in niche areas like cybersecurity. The stealth update to April’s preview version came with a brief statement on DeepSeek’s official website on Wednesday, noting that the model offered “significantly enhanced agent capabilities”, but...
- DeepSeek releases official V4 Pro model as it steps up expansion
DeepSeek releases official V4 Pro model as it steps up expansion Reuters