AI News Archive: August 14, 2026 — Part 1
Sourced from 500+ daily AI sources, scored by relevance.
- Google AI health coach to use Abbott glucose data
Abbott and Google are linking continuous glucose monitoring data with Google’s AI-powered health coaching tools, giving the Gemini-powered service access to another source of personal health information. Under a multiyear agreement, data from Abbott’s Lingo continuous glucose monitor will be integrated into the Google Health app. Users will be able to view glucose trends alongside […] The post Google AI health coach to use Abbott glucose data appeared first on AI News .
Score: 90🌐 MovesAug 14, 2026https://www.artificialintelligence-news.com/news/google-ai-health-coach-abbott-glucose-data/ - Waymo wins approval for drastic expansion in Los Angeles, Bay Area regions
Waymo announced today it received approval from California regulators to expand its service area across 18 counties, including the cities of San Diego and Sacramento.
- EXCLUSIVE: Apple trains its own AI model for China market with Alibaba's support, sources say
EXCLUSIVE: Apple trains its own AI model for China market with Alibaba's support, sources say Reuters
- Grok 4.6 x Cursor : Elon Musk Just Bought His Way Into the AI Coding War
Grok 4.6 does not dominate Claude or Codex. With frontier performance, aggressive pricing, and Cursor distribution, it may not need to. Continue reading on Towards AI »
- Anthropic Revenue Ahead of IPO Surges Over 14-Fold in Second Quarter
Anthropic PBC is telling prospective investors its second-quarter revenue jumped at least 14-fold versus the same period a year ago, according to documents seen by Bloomberg News.
- Google drops Gemini 3.7 Flash model, and it’s ready to handle your chores with the Spark agent
Google’s new Gemini 3.7 Flash model brings sizable gains in workflow automation and document handling, two areas that could make Spark much more useful.
- Nvidia scales back $250 billion OpenAI data center guarantee, WSJ reports
Nvidia scales back $250 billion OpenAI data center guarantee, WSJ reports Reuters
- OpenAI’s Chief Revenue Officer Is Leaving After 8 Months. She’s Just the Latest Executive to Head for the Exit
What is going on at OpenAI? The ongoing exodus of top talent has critics speculating ahead of an expected IPO.
- Samsung health AI models analyse wearable biosignal data
Samsung Research America’s Digital Health Team has presented two AI foundation models designed to learn from wearable biosignals. The work centres on data captured by smartwatches, including heart activity, sleep, and physical activity. The company discussed its Connected Care vision at the Health Forum during Galaxy Unpacked in July 2026. Samsung described a future of […] The post Samsung health AI models analyse wearable biosignal data appeared first on AI News .
Score: 82🌐 MovesAug 14, 2026https://www.artificialintelligence-news.com/news/samsung-health-ai-models-analyse-wearable-biosignal-data/ - Fractal Analytics beta launches India’s first healthcare AI model under India AI Mission
The pilot in partnership with the Brihanmumbai Municipal Corporation (BMC) will provide a Health AI Chatbot on WhatsApp
- Gemini 3.7 🤖, GPT-5.6 Sol Ultrafast ⚡, Anthropic $2T IPO 💰
Gemini 3.7 🤖, GPT-5.6 Sol Ultrafast ⚡, Anthropic $2T IPO 💰
- Databricks valuation hits $190bn after second $5bn raise of 2026
The data and AI company's valuation has jumped from $134bn to $190bn in just six months. Read more: Databricks valuation hits $190bn after second $5bn raise of 2026
Score: 80💰 MoneyAug 14, 2026https://www.siliconrepublic.com/business/databricks-valuation-hits-190bn-after-second-5bn-raise-in-2026 - OpenAI previews 14× faster GPT-5.6 Sol for real-time agents
OpenAI announces GPT-5.6 Sol, claiming 14x speed increase for real‑time agent applications.
Score: 80🤖 ModelsAug 14, 2026https://aibreakfast.beehiiv.com/p/openai-previews-14-faster-gpt-5-6-sol-for-real-time-agents - OpenAI on Track to Double Revenue Ahead of IPO
OpenAI is on track to generate more than $40 billion in annualized revenue, roughly double its run rate at the end of 2025, according to sources. Bloomberg’s Rachel Metz explains how growth in paying ChatGPT users, enterprise adoption and coding assistant Codex are driving the surge, and why the enormous cost of computing power remains a key challenge for the AI company. She joins Tim Stenovec on "Bloomberg Tech." (Source: Bloomberg)
Score: 80💰 MoneyAug 14, 2026https://www.bloomberg.com/news/videos/2026-08-14/openai-on-track-to-double-revenue-ahead-of-ipo-video - GPT-5.6 Sol goes 14x faster as OpenAI launches Ultrafast mode powered by Cerebras
OpenAI is launching "Ultrafast," a new inference mode that delivers GPT-5.6 Sol at up to 750 output tokens per second, powered by Cerebras hardware from their $10 billion partnership. Together with "Standard" and "Fast," Ultrafast creates a three-tier pricing structure that turns inference speed into its own product. The article GPT-5.6 Sol goes 14x faster as OpenAI launches Ultrafast mode powered by Cerebras appeared first on The Decoder .
Score: 80🌐 MovesAug 14, 2026https://the-decoder.com/gpt-5-6-sol-goes-14x-faster-as-openai-launches-ultrafast-mode-powered-by-cerebras/ - GLM-5.3 is here with advanced cyber capabilities — and reportedly already found a 'serious vulnerability' in Cursor
Chinese AI startup Z.ai, known internationally for its growing lineup of powerful, largely open source GLM series of language models, today released GLM-5.3 with substantial gains in long-horizon coding and a more consequential — and potentially sensitive — jump in cybersecurity capabilities. Already, GLM-5.3's cyber capabilities have found a "potentially serious vulnerability in Cursor," the AI coding startup recently acquired by SpaceX , according to z.ai developer advocate Lou, posting on X . VentureBeat also tagged Cursor for confirmation on X and is awaiting response. GLM-5.3 is available initially only through the company's GLM Coding Plan and ZCode coding environment, while API access and open weights are coming later, "once safety evaluation and hardening are complete," according to the company. Z.ai says it plans to release weights approximately two weeks after launch. For enterprise developers, the notable part of the release is not simply another round of benchmark improvements. Z.ai says GLM-5.3 uses the same base model as GLM-5.2, with the improvements coming entirely from scaling post-training across more environments, more diverse tasks and additional reinforcement-learning compute. That makes GLM-5.3 something of a test of how far a frontier-scale base model can be pushed without another expensive pretraining cycle. “Scaling post-training is all we did for GLM-5.3,” Z.ai wrote in its technical announcement. The results suggest considerable headroom. But they have also produced an unusual problem for an open-model developer: according to Z.ai, cybersecurity capabilities improved faster than anticipated as training scaled, particularly as tasks progressed from vulnerability identification toward constructing complete exploitation chains. Reuters reported Friday that Z.ai is also introducing controls around some of the model's more advanced capabilities, including a “trusted access” approach for sensitive functionality. A large jump in coding without another base model GLM-5.3 builds on the 743-billion-parameter-scale base model behind GLM-5.2 rather than replacing it. Z.ai instead expanded the post-training system it had already assembled around long-horizon reinforcement learning. Those environments increasingly resemble complete engineering jobs rather than isolated programming exercises. Z.ai describes scenarios in which an agent receives access to codebases, documentation, compute clusters, storage systems and experimental results, then has to diagnose problems, modify systems, run experiments and demonstrate a measurable improvement while preserving correctness. Some tasks are designed to approximate several days of work for an experienced engineer. The approach produced sizable generation-over-generation improvements on Z.ai's reported evaluations. GLM-5.3 jumps from 4.6 to 28.3 on Terminal-Bench 3.0, from 46.2 to 66.9 on DeepSWE v1.1, and from 26.2 to 48.2 on AutomationBench. On Agents' Last Exam CLI, it improves from 23.8 to 28.5. The model does not dominate every frontier competitor. Z.ai's own benchmark table shows GPT-5.6 Sol at 34.6 and Claude Fable 5 at 33.7 on Terminal-Bench 3.0, compared with GLM-5.3's 28.3. On DeepSWE v1.1, GLM-5.3 scores 66.9, compared with 72.7 for GPT-5.6 Sol and 69.7 for Fable 5. But Z.ai is also emphasizing efficiency rather than benchmark position alone. On its private Z.ai Code Bench, GLM-5.3 reaches a 34.5% result at its Max reasoning setting while consuming roughly 75,000 output tokens per task. GLM-5.2 reaches 23.4% while consuming approximately 96,000. At High effort, GLM-5.3 reaches 31.4% at roughly 50,000 output tokens, compared with Z.ai's reported 29.5% for Claude Opus 4.8 using 120,000. Because Code Bench is Z.ai's own private evaluation, those comparisons should be treated as company-reported results rather than independent measurements. Still, reducing token consumption while improving task completion is operationally important for enterprises deploying coding agents, where long-running loops can make inference cost and latency compound quickly. Cyber capabilities developed faster than Z.ai expected The more unusual development is cybersecurity. Z.ai introduced vulnerability-discovery environments into GLM-5.3's post-training mix expecting the model to improve at finding software flaws. Instead, the company says capability began progressing further along the exploitation chain. “As we scaled post-training, cyber capability developed faster than we expected,” Z.ai wrote. On CyberGym, which tests vulnerability discovery and validation against source code, GLM-5.3 scores 84.5%, compared with 77.2% for GLM-5.2. That also edges Z.ai's reported scores for GPT-5.6 Sol at 83.6% and Mythos 5 at 83.8%. The advantage does not extend across the entire exploitation stack. GLM-5.3 scores 54.4% on ExploitBench, more than twice GLM-5.2's 24.4%, but remains well behind the 76.5% Z.ai reports for GPT-5.6 Sol and 78% for Mythos 5. Similarly, on ExploitGym, GLM-5.3 completes 105 tasks under a normalized two-hour budget and 130 under six hours, up from 29 and 39 for GLM-5.2. Fable 5 reaches 181 and 247, while GPT-5.6 Sol reaches 216 and 293. The direction of travel may matter more than the leaderboard position. Z.ai says work with security teams in China has resulted in 2,436 vulnerability findings across 269 projects after expert review, screening and deduplication. Its disclosure ledger lists 1,097 as critical or high severity, with 53 publicly disclosed and 2,383 still under embargo at the time of the release. That creates a tension increasingly facing frontier model providers: the same long-horizon agent capabilities that make models more useful for software engineering can also make them more capable security researchers — and potentially more capable offensive operators. GLM-5.3 also requires developers to change how they call the model Developers migrating existing GLM applications should pay attention to a breaking API behavior. GLM-5.3 supports three reasoning-effort levels — low , high and max — with max the default and Z.ai's recommended setting for coding. But unlike previous releases, thinking cannot be disabled. Applications currently sending thinking.type: "disabled" must change the value to enabled and specify a reasoning effort before switching the model identifier to GLM-5.3. Otherwise, Z.ai says the request will fail. That makes GLM-5.3 an actual migration rather than simply a model-name substitution for some production applications. From GLM-4.5 to GLM-5.3: Z.ai's rapid push into agentic engineering GLM-5.3 is the latest step in a rapid shift by Z.ai — formerly known as Zhipu AI — toward coding agents and long-running autonomous engineering workloads. GLM-4.5, released in July 2025, established much of that direction. The 355-billion-parameter mixture-of-experts model was designed to combine reasoning, coding and agent capabilities, while the smaller GLM-4.5-Air offered 106 billion total parameters. Z.ai released the models with open weights and emphasized integration with agent frameworks. GLM-4.6 followed in September, expanding context from 128,000 to 200,000 tokens and targeting coding, tool use and agent workflows in environments including Claude Code, Cline, Roo Code and Kilo Code. Z.ai also began placing greater emphasis on token efficiency in real-world coding evaluations rather than benchmark performance alone. The larger architectural jump came with GLM-5 in February 2026. Z.ai scaled the model from GLM-4.5's 355 billion parameters to 744 billion, with 40 billion active parameters, and increased pretraining data to 28.5 trillion tokens. It also introduced its “slime” asynchronous reinforcement-learning infrastructure and explicitly repositioned the GLM family around “agentic engineering” and long-horizon tasks. By June, GLM-5.2 had turned that strategy into a more direct enterprise proposition. The 753-billion-parameter model arrived with a stable 1-million-token context window, open weights under an MIT license and support across more than 20 coding environments. It also introduced IndexShare, which reuses an indexer across sparse-attention layers to reduce the computational burden of very long contexts. GLM-5.2 was priced at $1.40 per million API input tokens and $4.40 per million output tokens, with cached input priced substantially lower, positioning Z.ai as both a technical and pricing competitor to proprietary frontier labs. Z.ai's ambitions have been expanding outside model development as well. Reuters reported last month that Zhipu AI raised roughly HK$31.4 billion, or about $4 billion, through a Hong Kong share sale, with proceeds intended for areas including research and development, computing infrastructure, talent and business expansion. Taken together, the releases show a consistent progression: GLM-4.5 unified reasoning, coding and agents; GLM-5 substantially scaled the foundation model; GLM-5.2 attacked long-context and long-horizon engineering; and GLM-5.3 is now attempting to extract substantially more capability from that same foundation through post-training. Pricing, ZCode and availability GLM-5.3 is available now through Z.ai's GLM Coding Plan and ZCode. ZCode is the company's own coding-agent environment and supports long-running “Goal” tasks that plan, implement, test and verify work. It also offers remote control of running tasks and is available on macOS, Windows and Linux. Individual GLM Coding Plans currently start at a listed promotional price of $12.60 per month for Lite with 10,000 credits per week. Pro is listed at $56 per month with six times Lite usage, while Max costs $117.60 per month with 14 times Lite usage. Team Standard and Premium seats are listed at $88 and $188 per user per month, respectively. Z.ai has also moved the Coding Plan to a points-based quota system that separately accounts for input, cached-input and output tokens. Calls outside the company's weekday peak period consume 50% of the normal points. The company has not yet provided general GLM-5.3 API pricing in the supplied launch materials, making total production API cost difficult to compare directly with GLM-5.2 or competing frontier models until staged API access arrives. That staged release may ultimately be the most important part of GLM-5.3. Z.ai spent the past year pushing an open-model strategy centered on permissive weights, low-cost inference and compatibility with existing coding-agent ecosystems. GLM-5.3 demonstrates what happens when that strategy succeeds perhaps too well in one sensitive domain: better autonomous engineering also means better autonomous security research. The result is a model that advances Z.ai's coding ambitions while forcing the company to confront the same capability-versus-access tradeoff facing the largest closed frontier labs. For enterprise developers, GLM-5.3 is therefore worth watching for two reasons. Its coding results provide another indication that increasingly capable agents can emerge from better post-training and environments without continuously rebuilding the underlying foundation model. Its cybersecurity results show why deciding how those agents are distributed may become just as important as deciding how they are trained.
- AI infrastructure rises from niche to global asset class with $500bn Nvidia deal
AI infrastructure rises from niche to global asset class with $500bn Nvidia deal thenationalnews.com
- SpaceXAI completes its Cursor acquisition following Grok Bot and Grok 4.6 release
When SpaceX isn’t landing rockets, it’s apparently landing AI company deals. In February, the firm behind Starlink absorbed xAI , which includes Twitter-turned-X. In April, SpaceX inked a deal with Cursor, a competitor to Claude Code and OpenAI’s Codex. As of August 14, the purchase has been completed. SpaceXAI, as it’s now called, also recently released Grok Bot and Grok 4.6 models with the help of Cursor.
Score: 80💰 MoneyAug 14, 2026https://9to5mac.com/2026/08/14/spacex-lands-deal-to-likely-purchase-claude-code-and-openai-codex-competitor/ - Alibaba's Qwen team releases Qwen 3.8 models with open weights under the Apache 2.0 license
Alibaba's AI team Qwen has released new open model weights under the Apache 2.0 license with Qwen 3.8. The dense 27-billion-parameter model is designed to outperform the larger Qwen 3.7 Plus in coding and office tasks and natively processes up to 262,000 tokens of context. With this release, Qwen is targeting developers building local and agent-based applications. The article Alibaba's Qwen team releases Qwen 3.8 models with open weights under the Apache 2.0 license appeared first on The Decoder .
Score: 80🤖 ModelsAug 14, 2026https://the-decoder.com/alibabas-qwen-team-releases-qwen-3-8-models-with-open-weights-under-the-apache-2-0-license/ - Private credit roundup: Nvidia's half trillion for chips financing, plus others
Private credit roundup: Nvidia's half trillion for chips financing, plus others Reuters
- OpenAI talent exodus raises 'huge red flag' ahead of IPO
OpenAI's C-suite turnover gives investors another reason for concern as the company pushes toward a mammoth IPO.
- SpaceX Completes $60 billion Cursor Acquisition
SpaceX Completes $60 billion Cursor Acquisition The Information
Score: 80💰 MoneyAug 14, 2026https://www.theinformation.com/briefings/spacex-completes-60-billion-cursor-acquisition - Nvidia discloses $21 billion stake in SpaceX at end of second quarter
Nvidia's stake in SpaceX, which came through an investment in xAI, was worth about $21 billion at the end of the second quarter.
Score: 80💰 MoneyAug 14, 2026https://www.cnbc.com/2026/08/14/nvidia-discloses-21-billion-stake-in-spacex-at-end-of-second-quarter.html - Unitree set for Shanghai debut after $905m IPO
The company also shipped more than 5,500 humanoid robots, compared with 5,168 sold by Shanghai-based rival Agibot.
Score: 80💰 MoneyAug 14, 2026https://www.techinasia.com/unitree-outships-us-humanoid-rivals-as-china-leads-hardware - Google Gemini 3.7 Flash Launched Globally With Enhanced Capabilitiesfor For Web Development and Software Engineering
Google has launched its new AI model, namely Gemini 3.7 Flash, which succeeds the July-released Gemini 3.6 Flash model. Google claims that the Gemini 3.7 Flash offers enhanced capabilities over its predecessor in multiple departments. The AI model is claimed to provide improved intelligence for complex workflows, coding, workflow automation, document comprehension, we...
- Nvidia’s $70 Bn AI Bet May Face a $2 Tn Reality Check as Anthropic Eyes IPO
Nvidia’s $70 Bn AI Bet May Face a $2 Tn Reality Check as Anthropic Eyes IPO apac.entrepreneur.com
- OpenAI’s Explosive Growth Continues | Bloomberg Tech 8/14/2026
Bloomberg’s Tim Stenovec takes a look at OpenAI, which is on track to generate annualized revenue of more than $40 billion, roughly doubling its run rate from the end of 2025. Plus, how a New York City bill backed by Mayor Zohran Mamdani could force Amazon to directly employ its delivery workers, and Meta's former chief AI scientist joins a new venture firm to invest in early AI startups. (Source: Bloomberg)
Score: 80💰 MoneyAug 14, 2026https://www.bloomberg.com/news/videos/2026-08-14/bloomberg-tech-8-14-2026-video - Alibaba releases Qwen3.8-Max under custom license
The release is Alibaba’s first move from a closed application programming interface model to open weights for a Max model.
Score: 80🤖 ModelsAug 14, 2026https://www.techinasia.com/alibaba-releases-qwenimage30-complex-layouts - Inside Kimi K3: How Moonshot AI Built the Largest Open-Source Model
Moonshot AI recently released Kimi K3, the largest open-weight model available as of mid-2026, and the first open model to reach the 3-trillion-parameter class. While most labs chase scale by simply adding more hardware, Moonshot took a different path. Instead of just building a bigger model, they redesigned how the model remembers things. That redesign is really the whole story of K3. The Road to K3 The story starts with Kimi K1.5 , which focused on scaling reinforcement learning and improving the model’s basic reasoning ability. By mid-2025, the team released Kimi K2 , a roughly 1-trillion-parameter model that pushed further on architecture and training pipelines. Not long after, Kimi K2.5 added native multimodal support and stronger agentic skills, meaning the ability to use tools and carry out multi-step tasks with less human guidance. As these models grew, they ran into a bottleneck that every large language model eventually hits: memory. The Memory Problem Standard transformer models remember everything. Every time the model generates a new word, it looks back at every previous word in the conversation. This lookup relies on something called the KV cache, which stores a representation of every token seen so far. The problem is that this cache grows with the conversation. Storage needed for it scales linearly with context length, but the total computation needed to keep re-checking that growing cache scales quadratically. In plain terms, double the conversation length and you don’t just double the cost, you roughly quadruple it. At the scale of a multi-trillion-parameter model with a million-token context window, that cost becomes something only the largest labs can absorb. To build a genuinely massive model without that cost exploding, Moonshot needed the model to forget selectively, rather than hoarding every detail forever. Kimi Linear: The Testbed In October 2025, Moonshot released Kimi Linear , a smaller 48-billion-parameter model built to test this idea. This is where they developed Kimi Delta Attention (KDA) , the mechanism that would later be scaled up to power K3. Here’s a simple way to picture the difference between standard attention and KDA: Standard attention is like a notebook. Every new piece of information gets written on a fresh page, and the notebook keeps growing forever. KDA is like a whiteboard with a fixed amount of space. When new information arrives, nothing new gets added. Instead, the model checks what has changed and updates the board in place. Using this approach, Kimi Linear cut memory use by around 75% compared to standard attention, with only a small trade-off in raw capability. The Delta Rule: How KDA Actually Forgets The mechanism behind KDA is called the Delta Rule , and it is what lets the model manage a fixed-size memory instead of a constantly growing one. Instead of stacking new entries on top of old ones, the Delta Rule updates memory in place. If a meeting moves from Tuesday to Thursday, the model does not keep both facts around. It erases “Tuesday” and writes “Thursday” into the same memory slot. Two learned values control this process for each attention head: β (writing strength) : controls how strongly new information overwrites what is already stored. α (forgetting factor) : controls how much of the old state is kept versus allowed to fade. KDA takes this further by giving each attention head 128 separate forgetting dials, rather than one shared setting. This lets the model be very selective. It can drop small talk and filler almost immediately, while holding on tightly to a specific name or date from hundreds of thousands of tokens earlier in the conversation. This “erase then write” approach also prevents new information from interfering with older memories, which was one of the main weaknesses of earlier linear attention designs. KDA itself builds on earlier research into Gated DeltaNet, which first introduced the idea of learning how much information to write into a recurrent memory state. Kimi K3’s Core Architecture Kimi K3 is a 2.8-trillion-parameter model, currently the largest open-weight model released, and Moonshot describes it as the first open model in the 3-trillion-parameter class. It supports a 1-million-token context window and understands text, images, and video within a single model, using a vision encoder called MoonViT-V2 (about 401 million parameters) to handle image and video input, alongside a roughly 160,000-token vocabulary. A few architectural pieces work together to make this possible. Kimi Delta Attention (KDA), for most layers. According to Moonshot’s technical report, K3 has 93 attention layers in total, and 69 of them use KDA to keep memory use flat regardless of context length. Moonshot reports that this delivers noticeably faster decoding at million-token context lengths compared to standard attention. Gated Multi-Head Latent Attention (Gated MLA), for the rest. KDA alone risks losing exact details over very long stretches, since it is built for efficient forgetting rather than perfect recall. So the remaining 24 layers use Gated MLA instead, an evolution of the Multi-Head Latent Attention architecture originally introduced by DeepSeek, which compresses the KV cache into a smaller latent representation. In K3, these layers are interleaved with the KDA layers at roughly a 3-to-1 ratio: the KDA layers handle efficient, fixed-size memory, while the Gated MLA layers act as a compressed full-recall path that preserves exact, token-by-token history for details that cannot be allowed to fade. Attention Residuals (AttnRes). In deep models, information from early layers can get diluted as it passes through many later layers, sometimes called PreNorm dilution. AttnRes acts as a drop-in replacement for standard residual connections, letting later layers selectively pull in representations from earlier layers instead of relying only on a single accumulated stream. Moonshot reports this delivers around 25% better training efficiency for well under 2% additional compute cost. Stable LatentMoE. K3 uses a Mixture-of-Experts design with 896 experts, of which only 16 are activated per token. Out of the 2.8 trillion total parameters, Moonshot’s technical report states that roughly 104 billion are active for any given token. This sparsity is what allows the model to reach such a large total parameter count while keeping the actual compute per token manageable. Put together, Moonshot says these architectural changes, combined with updates to training and data, give K3 roughly 2.5x the overall scaling efficiency of K2, meaning it converts a given amount of compute into more usable capability than its predecessor. Training and Optimization Details Running stable training at 2.8 trillion parameters with only 16 of 896 experts active required a few extra tricks beyond the core attention design. Quantile Balancing. With so few experts active per token, keeping the workload evenly spread across all 896 experts is a real challenge. Instead of the usual approach of adding an auxiliary loss term to nudge routing toward balance, Quantile Balancing sets each expert’s routing bias directly from router-score quantiles, so load balancing falls out of the routing math itself rather than needing a separate, sensitive hyperparameter to tune. Per-Head Muon. Muon is a training optimizer that has become popular for large-scale model training. Moonshot extended it so that attention heads are optimized independently rather than as one block, which the team says gives more adaptive learning at this scale. Sigmoid Tanh Unit (SiTU). This is a custom activation function used in place of more common choices like GeLU or SwiGLU, intended to give the model finer control over activations. Quantization-aware training. Rather than training at full precision and quantizing afterward, K3 is trained with quantization awareness starting from the supervised fine-tuning stage. The released weights use the MXFP4 format with MXFP8 activations, which keeps the model runnable on a wider range of hardware without a separate, lossy quantization step after training. Benchmarks and Real-World Performance At launch on July 16, 2026, Kimi K3 took the top spot on Arena’s WebDev leaderboard (a human blind-vote benchmark for AI-generated front-end code) with a score of 1,679, ahead of Claude Fable 5 and GPT-5.6 Sol at the time. It also led on several sustained coding and agentic benchmarks, including Program Bench, SWE Marathon, BrowseComp, and OmniDocBench. Moonshot has been direct that K3 does not lead across the board. The company’s own materials describe it as trailing the strongest proprietary models, including Claude Fable 5 and GPT-5.6 Sol, on overall performance and on benchmarks like FrontierSWE and HLE-Full, even while outperforming other tested models on its evaluation suite. As is typical with vendor-reported benchmarks, results depend heavily on the exact test harness, reasoning settings, and context management used, so these numbers are best read as a general signal of strength rather than an exact ranking. Release Timeline and Background Kimi K3 comes from Moonshot AI, a Beijing-based company founded in 2023 by Yang Zhilin, who studied computer science at Tsinghua University, completed a machine learning PhD at Carnegie Mellon, and worked on earlier long-context research such as Transformer-XL and XLNet before starting Moonshot. Long-context modeling has been a consistent focus for the company since its earliest products, and KDA is best understood as the latest step in that same line of work. K3 was announced on July 16, 2026, with the full open-weight release following on July 27, 2026, distributed as roughly 96 shards totaling around 1.56 terabytes, under a custom Kimi K3 License. It is available through Kimi.com, Kimi Work, Kimi Code, and the Kimi API, and Moonshot has also contributed a KDA implementation to the vLLM community to support self-hosted deployment. Why This Matters The throughline across K1.5, K2, K2.5, Kimi Linear, and now K3 is a shift away from “just add more compute” and toward teaching models to manage memory the way a person might: keep the important stuff precise, let the unimportant stuff fade, and avoid paying to re-read the entire conversation from scratch every time something new is said. That combination of selective memory (KDA), targeted exact recall (Gated MLA), better information flow across depth (AttnRes), and aggressive but carefully balanced sparsity (Stable LatentMoE) is what let Moonshot push past the 2-trillion-parameter mark on open weights while keeping the model usable at a million tokens of context. Inside Kimi K3: How Moonshot AI Built the Largest Open-Source Model was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.
- Anthropic’s Biggest Acquisition Could Be a $6 Billion Bet on Making Claude Run Faster
Decart’s optimization technology aims to squeeze more computing power from existing hardware—a potentially enormous prize as AI infrastructure spending soars.
- Lovable raises $400M, Duolingo acquires Animade, and Isembard is a politician's "wet dream"
This week, we tracked more than 35 tech funding deals worth over €763 million and over 10 exits, M&A transactions, rumours, and related news stories across Europe.If email is more your thing, you ...
- WRITER Makes Agentic AI Economically Sustainable at Enterprise Scale With Palmyra X6 Release and Major Harness Upgrades
SAN FRANCISCO — August 13, 2026 — WRITER, the enterprise AI agent platform trusted by the world’s leading Fortune 500 brands, today released Palmyra X6, its new flagship model that brings frontier-competitive performance to the critical workflows marketing and revenue teams run every day work,... The post WRITER Makes Agentic AI Economically Sustainable at Enterprise Scale With Palmyra X6 Release and Major Harness Upgrades appeared first on WRITER .
- OpenAI annual revenue set to top $40 billion
The figures, reported by Bloomberg, mark one of the quickest growth stories in history, as the company prepares for what is expected to be a blockbuster IPO in the coming months.
Score: 78🌐 MovesAug 14, 2026https://www.semafor.com/article/08/14/2026/openai-revenue-set-to-top-40-billion - Micron Ventures Launches $250 Million Fund to Invest in the Next Generation of AI
Micron Ventures Launches $250 Million Fund to Invest in the Next Generation of AI markets.businessinsider.com
- Anthropic Revenue Jumped 14 Times in Second Quarter
Anthropic Revenue Jumped 14 Times in Second Quarter The Information
Score: 78🌐 MovesAug 14, 2026https://www.theinformation.com/briefings/anthropic-revenue-jumped-14-times-second-quarter - EXCLUSIVE: US to tell partners they must pick sides in AI race with China
EXCLUSIVE: US to tell partners they must pick sides in AI race with China Reuters
Score: 78🌐 MovesAug 14, 2026https://www.reuters.com/world/china/us-tell-partners-they-must-pick-sides-ai-race-with-china-2026-08-14/ - Launch of DeepSeek’s Harness marks its strategic pivot towards autonomous agentic AI
Chinese artificial intelligence company DeepSeek is venturing into a new battleground beyond large language models, launching a developer preview of its long-anticipated Harness – a software framework that helps developers turn AI models into autonomous agents. The release on Thursday marks a strategic pivot as DeepSeek moves to build foundational digital scaffolding for AI agents – systems capable of using AI models to operate external software, run code, and complete complex jobs on their...
- DeepSeek launches V4 Pro model with enhanced AI agent capabilities
Chinese artificial intelligence company DeepSeek officially launched the latest version of its flagship V4 Pro model in the early hours of Wednesday, further enhancing its capabilities in AI agents and software engineering.
- Nvidia Downsizes Plans for $250 Billion Guarantee of OpenAI Data Center
Investors are worried about the chipmaker’s risk exposure as it wields its balance sheet to bolster demand for its AI chips.
- Anthropic details unreleased Model 2, new alignment concerns in latest AI risk report
Anthropic PBC today revealed that it has developed an artificial intelligence model more capable than Claude Mythos 5. The company detailed the algorithm in the latest edition of its AI alignment report. The document, which is published every three to six months, outlines the potential risks posed by the company’s large language models. The newest […] The post Anthropic details unreleased Model 2, new alignment concerns in latest AI risk report appeared first on SiliconANGLE .
- Apple Intelligence in China: Alibaba Backs a Custom AI Model
Apple reportedly trained a China-specific AI model with Alibaba, changing how Apple Intelligence may work for users and businesses in mainland China. The post Apple Intelligence in China: Alibaba Backs a Custom AI Model appeared first on TechRepublic .
Score: 76🤖 ModelsAug 14, 2026https://www.techrepublic.com/article/news-apple-china-ai-model-alibaba-intelligence-apac/ - Goldman in talks with investors on Nvidia financing deal after landing prized role, sources say
Goldman in talks with investors on Nvidia financing deal after landing prized role, sources say Reuters
- Berlin’s AI InsurTech startup omni:us acquired by Dortmund’s adesso to embed AI into core insurance systems
adesso, a Dortmund-based provider of insurance-specific software and technology solutions, has acquired omni:us, a Berlin-based startup specialising in AI-powered claims processing. The transaction combines adesso’s insurance and software expertise with omni:us’ AI technology. adesso also plans to progressively integrate omni:us technology into its core insurance platform, the in|sure Ecosphere. omni:us solutions will be further developed […] The post Berlin’s AI InsurTech startup omni:us acquired by Dortmund’s adesso to embed AI into core insurance systems appeared first on EU-Startups .
- Nvidia’s $500 billion plan envelops Wall Street in its AI frenzy
Nvidia has unveiled a major financing initiative aimed at supporting artificial intelligence chip acquisitions. By collaborating with prominent financial institutions, the plan seeks to provide substantial debt solutions tailored for AI startups. This multifaceted financing strategy is designed to address the skyrocketing demand for computing power, ultimately fostering the essential digital and AI infrastructure development needed in the tech industry.
- BYD-Backed Deep-Sea Robot Maker Shenhai Zhiren Raises 500 Million Yuan, Debuts Underwater Agent Model SEAgent 1.0
Guangzhou deep-sea robot company Shenhai Zhiren closed a 500 million yuan A round with new investors including GGV Capital, Cathay Capital's TotalEnergies-backed fund, and others, while existing backers BYD, Lightspeed China, and GGV followed on. Its Phoenix 600 cable-burial robot, operating at 3,000 meters, won a 100 million yuan order from a UAE telecom group.
Score: 74🤖 ModelsAug 14, 2026https://pandaily.com/shenhai-zhiren-500-million-a-round-deep-sea-robots-byd-backed-seagent-aug2026 - China's Pony.ai, Uber to jointly deploy over 2,000 robotaxis in Europe
China's Pony.ai, Uber to jointly deploy over 2,000 robotaxis in Europe Reuters
Score: 73🌐 MovesAug 14, 2026https://www.reuters.com/technology/chinas-ponyai-uber-jointly-deploy-over-2000-robotaxis-europe-2026-08-14/ - Chipmaker CXMT overtakes Tencent as most valuable Chinese firm amid AI frenzy
Chipmaker CXMT overtakes Tencent as most valuable Chinese firm amid AI frenzy The Straits Times
- Daily Digest: Workday may be in play, Databricks raises more than expected
San Francisco officials approved the first stage of a massive housing development despite opposition from Mayor Daniel Lurie and local neighborhood representatives.
- OpenAI CFO Friar tells investors that enterprise business now bigger than consumer by revenue
OpenAI CFO Sarah Friar met with investors following a week of turmoil in the C-suite that included the sudden departure of revenue chief Denise Dresser.
Score: 73🌐 MovesAug 14, 2026https://www.cnbc.com/2026/08/14/openai-cfo-friar-tells-investors-that-enterprise-bigger-than-consumer.html - OpenAI and Anthropic in price war as Chinese AI rivals gain ground
US groups release cheaper models after new challenges to their trillion-dollar ambitions
Score: 72🌐 MovesAug 14, 2026https://www.ft.com/content/32a70a3c-7d28-40b4-808e-36edb58c7d01?syn-25a6b1a6=1