🤖 Models AI News
New AI model releases and updates. From GPT and Claude to open-source models, we score and categorize the top new AI model news.
- Google Deepmind's WeatherNext predicts cyclone tracks and intensity at the same time
Deepmind's new weather AI forecasts tropical cyclones about a day further ahead than leading operational models, matching a decade of progress in traditional weather forecasting. Code and model weights are open-source on GitHub. The article Google Deepmind's WeatherNext predicts cyclone tracks and intensity at the same time appeared first on The Decoder .
Score: 84🤖 ModelsAug 9, 2026https://the-decoder.com/google-deepminds-weathernext-predicts-cyclone-tracks-and-intensity-at-the-same-time/ - NVIDIA Releases NemotronLabs VoiceChat 11B: An Open Full-Duplex Speech-to-Speech Model with ~450 ms Turn-Taking and Live Tool Calling
NVIDIA Releases NemotronLabs VoiceChat 11B: An Open Full-Duplex Speech-to-Speech Model with ~450 ms Turn-Taking and Live Tool Calling MarkTechPost
- Google's DiffusionGemma proves you don't need to train from scratch to build a text diffusion model
Instead of training a new model from scratch, Google DeepMind retrofitted Gemma 4 into a diffusion model using less than 10 percent of the original training budget. DiffusionGemma generates 256 tokens in parallel instead of one at a time, hitting about 1,500 tokens per second. Quality still trails the original autoregressive model in benchmarks, especially on reasoning tasks. The article Google's DiffusionGemma proves you don't need to train from scratch to build a text diffusion model appeared first on The Decoder .
- AI Models and the Houdini Act: Kimi K3 and Meta’s Muse Spark 1.1 Are the Culprits Now
If Anthropic took showed its marketing chutzpah by claiming Claude Mythos was too dangerous for the real world, OpenAI outdid it by claiming their model had autonomously committed cyberattack. The trend is now assuming pandemic proportions as first Anthropic and now Kimi K3 (Moonshot AI) and Muse Spark 1.1 (Meta) have all reported their own […] The post AI Models and the Houdini Act: Kimi K3 and Meta’s Muse Spark 1.1 Are the Culprits Now appeared first on CXOToday.com .
- OpenAI flags its new Astra model as potentially reaching the highest cybersecurity risk level for the first time
Internal tests of OpenAI's new AI model Astra show cybersecurity capabilities so strong that the company can no longer rule out the highest risk level in its own safety framework. Parts of Astra's development have been paused. The move follows recently disclosed incidents in which autonomous AI agents infiltrated OpenAI's own infrastructure undetected for weeks. The article OpenAI flags its new Astra model as potentially reaching the highest cybersecurity risk level for the first time appeared first on The Decoder .
- AI Model Invents Completely New Viruses Unknown to Nature
What could possibly go wrong? The post AI Model Invents Completely New Viruses Unknown to Nature appeared first on Futurism .
- China's Largest AI Model Is Being Developed at Bytedance
Bytedance is training an AI model with up to ten trillion parameters, according to the Financial Times. That's three times the size of Moonshot's Kimi K3, currently the largest Chinese model. The article China's Largest AI Model Is Being Developed at Bytedance appeared first on The Decoder .
Score: 78🤖 ModelsAug 7, 2026https://the-decoder.com/chinas-largest-ai-model-is-being-developed-at-bytedance/ - Chinese startup Moonshot's AI model breaks out of testing environment, researchers say
Chinese startup Moonshot's AI model breaks out of testing environment, researchers say Reuters
- EXCLUSIVE: Alibaba plans to charge big users of its next open-source AI model, sources say
EXCLUSIVE: Alibaba plans to charge big users of its next open-source AI model, sources say Reuters
- Allstate Introduces Large Language Model, ALLIE
Allstate CEO Tom Wilson said the insurer has a “technology-drive strategy, not a strategy supported by technology.” And the next step in the strategy is the build of ALLIE – Allstate’s Large Language Intelligent Ecosystem. Speaking to analysts during a …
- CASIA's PhiZero Gives World Models a 'Physical Language' and Cuts Tokens 175x
Researchers at the Institute of Automation under the Chinese Academy of Sciences (CASIA) published PhiZero, a world model that reasons in a learned discrete 'physical language' before rendering video. The approach collapses the number of tokens needed to represent a four-second clip by a factor of about 175.
Score: 51🤖 ModelsAug 7, 2026https://pandaily.com/casia-phizero-world-model-physical-language-175x-token-reduction-aug2026 - New AI model reveals the volume of the world's glaciers
How much ice is stored in the world's glaciers? And where exactly is it located? A new study led by Ca' Foscari University of Venice, in collaboration with the Institute of Polar Sciences of the National Research Council of Italy (CNR-ISP), provides an updated global map of glacier ice volume.
- The next age of LLMs? Dev gets a small LLM running at 10 tokens a second locally on a $10 microcontroller
A 28.9M-parameter model runs on a microcontroller costing less than $10 at 9.88 tokens a second because 25M of those parameters never leave flash storage
- GPT-5.6 Luna default 🌙, Agent Plugins 🔌, AMD Taalas acquisition 🧩
GPT-5.6 Luna default 🌙, Agent Plugins 🔌, AMD Taalas acquisition 🧩
- Chinese AI Model Kimi K3 Escapes Sandbox in Third-Party Test, Researchers Say
Chinese firm Moonshot’s latest artificial intelligence model broke out of a cyber-testing environment, researchers said, in the latest incident that raises concerns about how well AI companies control their technology. Moonshot’s Kimi K3 was able to find its way out …
- OpenAI Pauses Some Work on New Astra Model on Cyber Concerns
OpenAI is pausing some internal work around one of its upcoming artificial intelligence models to implement stricter safeguards after the system was found to be significantly more adept at cybersecurity tasks.
- OpenAI puts the brakes on a new model because it’s supposedly too powerful
OpenAI says it is pausing "internal activities" around an in-development AI model, Astra, because it doesn't yet meet new security standards the company is putting in place. The announcement follows its recent disclosure that OpenAI models accidentally hacked Hugging Face. Anthropic and Meta have also since admitted that they had AI models that went rogue […]
- Chinese AI model breaks through constraints
The incident combes after other AI models went rogue, accessing the internet and hacking other companies, generating concerns about their abilities to bypass human controls.
- Did China’s Kimi K3 AI really find a way out of its safety sandbox?
Did China’s Kimi K3 AI really find a way out of its safety sandbox? YourStory.com
- China’s Kimi K3 AI model escapes isolated sandbox during security test: researchers
China’s top open-weight AI model Kimi K3 broke out of its isolated test environment during a cybersecurity evaluation, according to US security researchers, following similar high-profile incidents involving closed frontier models from OpenAI and Anthropic that highlight the growing challenge of constraining AI behaviour. Kimi K3, released last month by Beijing-based Moonshot AI, escaped from a supposedly isolated sandbox environment, accessed the open internet and found solutions on the...
- Chinese AI model Moonshot Kimi K3 also escaped its testing environment
Kimi K3 also found loopholes in its sandbox environment that allowed it to access the internet.
- AI models keep escaping their sandboxes, and Kimi K3 is the latest to join the party
Moonshot AI's Kimi K3 slipped past its sandbox during a security test, becoming the latest AI model caught wandering onto the open internet this summer.
- ByteDance targets mega AI model that could match Mythos scale, FT reports
ByteDance targets mega AI model that could match Mythos scale, FT reports Reuters
- Stanford Evo 2 AI model generates phages against E. coli
Stanford researchers have synthesised nearly 300 phages from DNA sequences produced by the Evo 2 generative AI model. Laboratory testing narrowed the group to 16 phages that showed particularly strong E. coli-killing activity. The work centres on bacteriophage ΦX174, pronounced “FYE-ex-1-7-4”. Brian Hie, an assistant professor of chemical engineering and Dieter Schwarz Foundation Stanford Data […] The post Stanford Evo 2 AI model generates phages against E. coli appeared first on AI News .
- A new AI model reveals the volume of the world's glaciers
A new AI model reveals the volume of the world's glaciers EurekAlert!
- No cloud, no GPUs, no problem: Liquid AI's new model LFM2.5-2.6B brings powerful AI agents to devices as small as a Raspberry Pi
Earlier this week, the AI startup Liquid, formed in 2023 by former MIT computer scientists, debuted LFM2.5-2.6B , a new open-weight language model designed specifically for agentic workloads. In release materials and a recent interview with VentureBeat, Liquid's researchers said LFM2.5-2.6B can run entirely on local hardware — from smartphones and laptops down to a Raspberry Pi — without relying on cloud inference or GPUs, unlocking edge AI applications and giving more options to enterprises working in regulated industries or with sensitive information they don't want to send up to the cloud. It's best suited for high-volume, well-defined agentic tasks that run locally — tool calling, document management, calendar and workflow automation, and always-on background routines — and for connectivity-limited environments like vehicles and robotics, though coding-heavy work is better left to larger models. Even for those businesses without such concerns, the appeal of running performant, task-specific agents at the cost of essentially electricity, may be enough to make the new model quite appealing. But the custom open weights license , as with Moonshot's larger frontier model Kimi K3 released last month, is worth a close look by enterprise legal teams. The basics LFM2.5-2.6B contains 2.6 billion parameters, supports a 128,000-token context window, and includes native tool calling. The somewhat tricky name is explained by the generation of model (2.5) combined with the parameter count (2.6B). Both the post-trained model and a base checkpoint (LFM2.5-2.6B-Base) for developers who want to fine-tune it are available now on Hugging Face , with day-one support for major inference stacks including llama.cpp, MLX, vLLM, SGLang, and ONNX — positioning it for deployment across consumer hardware, enterprise infrastructure, and embedded systems. Liquid also offers an open source fine-tuning framework, LEAP . Rather than positioning LFM2.5-2.6B as a competitor to the largest frontier models, the company is making a different argument: that a sufficiently capable small model can unlock categories of enterprise applications where latency, privacy, deployment flexibility, or inference costs matter more than absolute benchmark leadership. "I do also believe that the best models will be in the cloud, and there's no problem with that," Maxime Labonne, Liquid AI's head of post-training, told VentureBeat in an interview following the launch. "We want to make models for another type of user, and the best way of describing it is: you should use [edge AI] when you can't use a cloud model." Small enough for a Raspberry Pi Asked about the minimum viable hardware, Labonne said the model runs "very, very well" on CPUs — and that the LFM2 architecture underlying the model was explicitly designed around real-world CPU performance rather than GPU benchmarks. "I think the best example is a Raspberry Pi," he said. "We have a lot of demos that show that actually, it works pretty fast on the Raspberry Pi." Company-reported measurements indicate decoding throughput of approximately 220 tokens per second on an Apple M5 Max and 113 tokens per second on an AMD Ryzen AI Max+ 395, while using less than 2.5 GB of memory — and around 30 tokens per second on a smartphone. Users can try the models on their phones through Apollo, Liquid AI's mobile app. At the other end of the deployment spectrum, Liquid AI reports the model reaches nearly 15,000 output tokens per second on a single Nvidia H100 GPU under sustained concurrent load — roughly 1.3 billion tokens per day on one card. These figures are vendor benchmarks and have not been independently verified. For Labonne, memory footprint and speed are not conveniences but hard constraints that determine what can be deployed at all. "What we want to show is that it's a really good trade-off, because you get the level of quality that you get with much bigger models, but in a tiny, tiny form factor," he said. "You can deploy it in target devices where you are not able to deploy the other ones at all." Trained for agents instead of chatbots Liquid AI says LFM2.5-2.6B was developed around the assumption that language models are increasingly consumed through agent frameworks rather than traditional conversational interfaces. "Models are not consumed in chatbots anymore. They're really consumed through agentic harnesses, like OpenClaw, like Hermes Agent," Labonne said. "We wanted to make sure that this model is not just good at math or at code, but it's good at using tools." The model is pretrained on approximately 34 trillion tokens, with a vocabulary doubled to 128K to better support non-Latin scripts and a dedicated mid-training phase to extend the context window to 128K tokens for long-running agent workflows. Post-training follows a four-stage pipeline: supervised fine-tuning, teacher specialization (training separate expert models for domains like instruction following, math, code, and tool use), multi-domain on-policy distillation (MOPD) to merge those experts' capabilities back into a single student model, and finally agentic reinforcement learning. During that last stage, the model was trained directly inside production agent harnesses — including Hermes Agent and OpenClaw — on realistic productivity tasks involving research, coding, document management, tool invocation, and workflow automation, exposing it to those harnesses' actual tools, system prompts, and interaction patterns. Labonne described the pipeline overhaul as producing a "happy accident": gains that extended well beyond the agentic targets. "Through these new training techniques, we also got a lot better at everything. We got better at math, at instruction following. We've never been good at code, actually — and with this, we even got really good at code," he said. Building the model — and the harness Notably, Liquid AI also built its own agent harness rather than relying solely on existing frameworks, and demonstrated the model running inside it on a phone, planning and calling tools entirely on-device. "This is a harness running on a phone, and I don't know if there's any other harness running on a phone," Labonne said. The company had two reasons, he explained. The first was necessity — no phone-native harness existed. The second is a different interaction model: today's harnesses wait for a prompt, and Liquid AI wants assistants that act on their own. "We want proactive agents. We want agents that run in the background, check what you're doing, check your calendar, and based on this context, do tasks," he said. "That doesn't exist today, really." Co-designing the harness and model also lets the software compensate for the model's weak spots. "Everything that the model is bad at, the harness should help the model with — provide as much assistance as possible to make it more reliable," Labonne said. "End users don't care if it's the model or the harness. What they want is that the task is achieved at the end of the day." The model nevertheless works out of the box with established harnesses including Hermes Agent, OpenClaw, and Pi, served behind any OpenAI-compatible endpoint. Swap the harness, not the model For enterprise deployment, Labonne argued the release marks a shift in what small models can be used for. Until now, he said, local models made economic sense mainly as narrowly fine-tuned specialists — trained to do one thing at cloud-model quality, much faster and cheaper. Agentic capability changes that calculus, because the same model can be repurposed by changing the tools around it rather than the model itself. "You can have a calendar assistant, and you can reuse the same model and make a meeting assistant that will record what everybody said and summarize it — a bit like Granola, for example," he said. "You don't change the model; you just change the harness. You just change the tools around it. This gives much more generalizability, and it's a lot easier to do and a lot cheaper as well." He still recommends fine-tuning for production deployments whenever feasible: "If you don't fine-tune it, you leave some quality on the table. If you fine-tune it well, it's going to match the performance of GPT and Claude — really, if your task is not the most complex task in the world," he said, adding that the barrier to entry has collapsed: "The bar to be able to do fine-tuning now is super low. It's very accessible to everyone." How it stacks up against DeepSeek-V4-Flash, Google's Gemma and Alibaba's Qwen Liquid AI released its own benchmark comparison charts pitting LFM2.5-2.6B against the models enterprises are most likely to shortlist for the same edge deployments: Google's Gemma 4 E2B (5.1B parameters) and E4B (8B), and Alibaba's Qwen3.5-4B (4.7B) and Qwen3.5-9B (9.7B). A separate test by local AI client platform Atomic Chat found that LFM2.5-2.6B completed 35 tool calls to complete three tasks (checking weather and local time in six cities, converting one budget into six currencies, checking four hotels and booking for a date) 3.7 times faster than DeepSeek-V4-Flash (a whopping 284B parameters), the model has skyrocketed to the top of OpenRouter since its release last week. Gemma 4's small models are multimodal generalists, accepting image and audio input alongside text, and use a Per-Layer Embeddings design that keeps only a fraction of their weights active per token — which is why Google markets them by "effective" size (2.3B and 4.5B) despite total footprints of 5.1B and 8B. Alibaba's Qwen3.5 small series, released in March , is natively multimodal from 4B up and leans on scaled reinforcement learning to chase frontier-style reasoning — Alibaba touts the 9B model as matching or beating OpenAI's far larger gpt-oss-120B on reasoning benchmarks. LFM2.5-2.6B takes a narrower path: it is text-only, dense, and specialized for agentic work, with Liquid AI shipping separate vision and audio variants of the LFM family rather than folding everything into one checkpoint. Where Qwen's post-training reinforcement learning targets reasoning, Liquid's targets tool use inside real agent harnesses. The result, per the company's published numbers, is that the smallest model in the comparison leads every instruction-following benchmark (IFBench, Multi-IF, IFStruct) and nearly every tool-use benchmark — 77.83 on ToolSandbox versus 76.44 for Qwen3.5-9B, a model nearly four times its size — trailing only that 9B model on BFCLv4. On agentic evaluations it beats both Gemma models across the board and essentially ties the Qwens: 26.89 on BrowseComp+ versus 27.23 for Qwen3.5-9B. It also posts the best score on AA Omniscience, a knowledge benchmark that penalizes hallucination. The Qwen models keep the edge where their training focus lies: math (Qwen3.5-9B leads AIME25) and coding, where larger models retain an advantage on LiveCodeBench — though Labonne noted the gap is smaller than the parameter counts would suggest. "With LiveCodeBench v6, we might not be the best among these models, but we're also by far the smallest. Showing that we're competitive with them is already quite a big win for me," he said. One differentiator cuts the other way: licensing. Gemma 4 and Qwen3.5 ship under the permissive Apache 2.0 license — a change Google made specifically to court enterprises . DeepSeek-V4-Flash ships under a similarly permissive MIT License . Meanwhile, Liquid AI's revenue-gated license (detailed below) asks larger companies to strike a commercial deal. Enterprises above the threshold are effectively trading license friction for footprint and tool-use performance. Licensing reflects a commercial middle ground LFM2.5-2.6B is distributed under the LFM Open License v1.0, which permits use, modification, and redistribution — including commercial use — for organizations with less than $10 million in annual revenue. Commercial use by larger companies is not covered by the license, requiring a separate arrangement with Liquid AI; qualified nonprofits are exempt from the threshold for non-commercial and research purposes. Labonne framed the structure as a way to sustain model development — "the models are really the moats, so we need to be sensible in the way that we license them; otherwise, we cannot make money, so we can't make more models" — while characterizing the threshold as a light-touch mechanism in practice. Asked how the company would even know if a large enterprise quietly deployed the open weights, he was candid: "I think this is a question for our legal team, but personally, I don't know. And even if you're above $10 million, the only thing that we ask you is to contact us." The company pairs its licensed model releases with freely published research, he added, including new structured-output evaluations and a training technique that mitigates the repetition loops common in small models — a failure mode he noted Qwen models are "kind of guilty of." Small model, big enterprise implications The launch coincided with an announcement from MacPaw , the Ukrainian software company behind CleanMyMac and Setapp, of a long-term strategic partnership with Liquid AI to build an on-device AI stack for the Mac. Liquid AI will design and fine-tune foundation models for Eney, MacPaw's macOS assistant, running locally on Apple silicon through MacPaw's Elix inference engine and Mnemos memory layer, with results expected later this year. Labonne pointed to the deal as a concrete validation of the size argument: "One of the reasons why they chose us is also because the model is quite small, and they don't have all the memory budget to run the other models." The release arrives as hardware vendors, operating system developers, and enterprise software companies increasingly invest in local AI execution — and as agent harnesses proliferate across the industry. Liquid AI's bet is that deployment economics, not raw scale, will define an important segment of that market: agents running continuously, everywhere, at zero marginal token cost. Whether small, highly optimized agent models become a significant segment of enterprise AI will ultimately depend less on benchmark scores than on operational reliability. But Liquid AI's latest release suggests the next competitive frontier is no longer simply building larger models — it's building models small enough, and capable enough, to run wherever enterprise workflows already live.
- Claude Fable 5 AI finds a tiny formula that topples an 87-year-old math conjecture
A mathematician working at Anthropic says he used the AI model Claude Fable 5 to uncover a remarkably simple counterexample to the Jacobian conjecture, a famous problem that has resisted mathematicians for more than a century. The result shows that the conjecture is false in three dimensions and above, although the original two-dimensional version remains unsolved.
- GPT-5 turning one as OpenAI shares new Agent Plugins standard
OpenAI’s GPT-5 turns one tomorrow, and the company is marking the week by looking beyond individual models. Today, OpenAI introduced Agent Plugins, an open standard meant to let reusable AI-agent extensions work across compatible products.
Score: 83🤖 ModelsAug 6, 2026https://9to5mac.com/2026/08/06/gpt-5-turning-one-as-openai-shares-new-agent-plugins-standard/ - Qwen 3.8 Max’s Incredible Debut: Better and Cheaper
A model ran unsupervised for ten straight days, building and testing its own software harness. Continue reading on Towards AI »
- Alibaba’s latest AI model puts it back in the great game
The ecommerce and cloud company hopes to recoup some of its lost shine with Monday’s release
Score: 81🤖 ModelsAug 6, 2026https://www.ft.com/content/391c5f14-4bf2-4d24-92f0-65b912574513?syn-25a6b1a6=1 - DeepSeek V4 Flash 0731 Outscores V4 Pro at 5x Lower Cost
DeepSeek V4 Flash 0731 Beats Its Own Pro Preview at $0.14 a Million Tokens That’s the Real Story Continue reading on Towards AI »
- Google open-sources an AI model it says can help with earlier hurricane warnings
WeatherNext can deliver a 15-day forecast predicting storms' track and intensity.
Score: 76🤖 ModelsAug 6, 2026https://www.engadget.com/2232141/google-open-source-ai-model-can-help-with-earlier-hurricane-warnings/ - Meta AI Model Accessed Internet, Hacked Outside Firm
Meta Platforms Inc. said one of its artificial intelligence models accessed the internet and hacked into an outside service’s systems during cybersecurity testing, following other recent incidents across the AI industry that have escalated concerns about companies’ control over their …
- OpenAI updating ChatGPT with a smarter GPT-5.6 Sol and unlimited free chats
ChatGPT has become so much more than just chat over the last year. To that end, OpenAI is removing limits on text chat. ChatGPT will also improve with a smarter version of GPT-5.6 Sol. Details below.
Score: 69🤖 ModelsAug 6, 2026https://9to5mac.com/2026/08/06/openai-updating-chatgpt-with-a-smarter-gpt-5-6-sol-and-unlimited-free-chats/ - Into the Omniverse: How Open World Models Push the Frontier of Physical AI
In July, NVIDIA joined more than 200 companies and organizations in signing “Open Weights and American AI Leadership,” an open letter arguing that AI leadership will be measured not by any single frontier model but by whether an open ecosystem reaches every sector.
- Westlake University's Yu Kaicheng Builds a Concept World Model and Speaks for the First Time
After a 100-million-yuan seed-plus-angel round, Awomo founder Yu Kaicheng makes his first public case for an implicit, concept-space world model that beats the data-driven scaling law he once helped write.
Score: 48🤖 ModelsAug 6, 2026https://pandaily.com/westlake-awomo-world-model-100m-yuan-seed-round-aug2026 - Free ChatGPT Users Get Unlimited Text Chats and GPT-5.6 Luna
OpenAI today said it is making GPT–5.6 Luna the default model for Free and Go ChatGPT users, giving them access to a newer, more capable model to replace GPT–5.5 Instant. The company is letting Free and Go users access unlimited text chats, with no text-based rate limits. Limits will still apply for file uploads, images, and other ChatGPT tools. There's also a new Think button coming that turns on higher reasoning. For Plus and Pro users, GPT–5.6 Sol is more reliable with facts and gives more focused answers. OpenAI added a slider to let users select how much thought ChatGPT puts into a response, plus the company tweaked the detail the model offers for each question to prevent it from giving extra information when it's not needed. OpenAI says GPT–5.6 Sol will avoid unnecessary formatting and will offer a "helpful correction when simply agreeing wouldn't be useful." It is also supposed to make fewer mistakes, especially when answers depend on dates, numbers, sources, rules, or assumptions. The Sol update is only for Chat, and the version for Work and Codex is unchanged. ChatGPT's Instant and Thinking experiences have a more consistent tone and behavior so it doesn't feel like switching to a model with a different tone or style. Plus and Pro users have access to the updated GPT–5.6 Sol and the slider in ChatGPT starting today. GPT–5.6 Luna is becoming the default model for Free and Go users this week, with unlimited text chats and the Think button coming next week. The changes apply to the ChatGPT app for iPhone, iPad , and Mac. Tags: ChatGPT , OpenAI This article, " Free ChatGPT Users Get Unlimited Text Chats and GPT-5.6 Luna " first appeared on MacRumors.com Discuss this article in our forums
- WeatherNext: AI model achieves breakthrough in forecasting cyclones
WeatherNext: AI model achieves breakthrough in forecasting cyclones
- Our WeatherNext 2 AI model demonstrated a massive leap forward in predicting cyclones.
Google DeepMind’s WeatherNext 2 shows state-of-the-art accuracy in cyclone prediction.
- Meta says its AI model hacked another company, adding to worries about bots going rogue
Meta says its AI model hacked another company, adding to worries about bots going rogue The Denver Post
- Alibaba releases Qwen 3.8 Max with 16-day autonomous execution capabilities
Alibaba unveils Qwen 3.8 Max, a large language model capable of 16-day autonomous execution, expanding AI-driven automation.
Score: 82🤖 ModelsAug 5, 2026https://aibreakfast.beehiiv.com/p/alibaba-releases-qwen-3-8-max-with-16-day-autonomous-execution-capabilities - NVIDIA Releases Alpamayo 2 Super: A 34B Open Vision-Language-Action Model for Robotaxis and Autonomous Driving Under OpenMDW-1.1
NVIDIA Releases Alpamayo 2 Super: A 34B Open Vision-Language-Action Model for Robotaxis and Autonomous Driving Under OpenMDW-1.1 MarkTechPost
Score: 77🤖 ModelsAug 5, 2026https://www.marktechpost.com/2026/08/05/nvidia-alpamayo-2-super-open-vla-model-autonomous-driving/ - How Kimi k3 Runs 2.8 Trillion Parameters on Consumer Hardware in 2026
Moonshot AI’s Kimi k3 claims 2.8 trillion parameters, yet inference runs on GPUs with 4GB VRAM. I tested the actual pipeline to find out… Continue reading on Towards AI »
- Black Forest Labs makes FLUX 3 Video generally available and claims it beats Seedance 2.0
Black Forest Labs has launched FLUX 3 Video, which generates Full HD clips up to 20 seconds long with native audio and lip-synced dialogue in more than 14 languages. It can also render typography directly in scenes. BFL's own Elo rankings put it ahead of Gemini Omni Flash and Seedance 2.0. The article Black Forest Labs makes FLUX 3 Video generally available and claims it beats Seedance 2.0 appeared first on The Decoder .
- Alibaba Launches Qwen3.8 Foundation Model and Qianwen Office Agent, Reseating Itself in the Enterprise AI Race
Alibaba unveiled the Qwen3.8 flagship with 2.4T total parameters and 95B activation, alongside Qianwen Office for enterprise AI workflows. The base-model-plus-office strategy leverages DingTalk enterprise moat against Tencent WorkBuddy and ByteDance Doubao.
- ByteDance launches SeedRealtime full-duplex audio-video model
ByteDance has launched SeedRealtime, a native audio-video full-duplex model that can continuously process audio, video and text streams while listening and responding in real time. The model is designed to support interactions in which it can watch, listen and speak simultaneously. ByteDance has rolled the model out in the Doubao app, moving the technology from […]
Score: 69🤖 ModelsAug 5, 2026https://technode.com/2026/08/05/bytedance-launches-seedrealtime-full-duplex-audio-video-model/ - Xiaomi open-sources embodied-AI foundation model Xiaomi-Robotics-1
Xiaomi has open-sourced its embodied-AI foundation model Xiaomi-Robotics-1, the company’s technology account announced on Aug. 5. The release covers the full process from real-robot post-training to model deployment and includes code for related benchmark evaluations. Xiaomi-Robotics-1 was pretrained on more than 100,000 hours of UMI data and post-trained on more than 10,000 hours of cross-embodiment […]
Score: 68🤖 ModelsAug 5, 2026https://technode.com/2026/08/05/xiaomi-open-sources-embodied-ai-foundation-model-xiaomi-robotics-1/ - Amazon open sources its enterprise safe OpenClaw-alike
An agentic side project reaches the real world.
Score: 65🤖 ModelsAug 5, 2026https://www.thestack.technology/amazon-open-sources-its-enterprise-safe-openclaw-alike/ - Mistral's open model Shieldstral matches much larger safety models at a fraction of the size
Mistral's new 3B Shieldstral model checks AI inputs and outputs for safety violations using natural language yes-or-no questions instead of fixed categories. It matches models seven times its size in some benchmarks. Operators can set their own criteria at runtime rather than rely on a third party's category system, and the model can run locally. The article Mistral's open model Shieldstral matches much larger safety models at a fraction of the size appeared first on The Decoder .
Score: 62🤖 ModelsAug 5, 2026https://the-decoder.com/mistrals-open-model-shieldstral-matches-much-larger-safety-models/ - Meet Shieldstral: Mistral's tiny AI model built to keep larger AIs safe
Meet Shieldstral: Mistral's tiny AI model built to keep larger AIs safe YourStory.com