AI News Archive: August 11, 2026 — Part 20
Sourced from 500+ daily AI sources, scored by relevance.
- IBM, Together AI ink $240 million deal for Nvidia-powered AI inference cluster
IBM, Together AI ink $240 million deal for Nvidia-powered AI inference cluster Reuters
- IBM bets $240m on cheap, open-source inference to take on the hyperscalers
A multi-year deal with the startup Together AI will put Nvidia Blackwell systems on IBM Cloud, and the wager is that enterprises now care more about the cost of running AI than the prestige of the model doing the running. IBM has decided that the money in artificial intelligence is no longer only in building […] This story continues at The Next Web
- Anthropic's planned mega-IPO faces investor skepticism over Chinese rivals and political headwinds
Anthropic is preparing an IPO for September or October, according to the Wall Street Journal, potentially the largest ever. During investor meetings, the company, valued at $965 billion, is fielding tough questions about Chinese competition, tensions with the Trump administration, and protests against data center construction. The company's IPO valuation will likely set the benchmark for how the entire AI industry gets valued. The article Anthropic's planned mega-IPO faces investor skepticism over Chinese rivals and political headwinds appeared first on The Decoder .
- The Video Production Stack Now Fits on One Desk: LTX-2.5 Launches as NVIDIA-Accelerated Open Weights World Model
The Video Production Stack Now Fits on One Desk: LTX-2.5 Launches as NVIDIA-Accelerated Open Weights World Model MarkTechPost
- Nvidia Releases New Open Model
Company says Nemotron 3.5 Lightning can help enterprises save on AI token costs
- Nvidia's open-weight Nemotron 3.5 Lightning prioritizes speed over maximum intelligence
Nvidia's Nemotron 3.5 Lightning is an open-weights model with just 3.6 billion active parameters that matches OpenAI's gpt-oss-120b on the Intelligence Index despite being four times smaller. At nearly 670 tokens per second, it's also the fastest model in the comparison, showing Nvidia is betting on efficiency over raw size. The article Nvidia's open-weight Nemotron 3.5 Lightning prioritizes speed over maximum intelligence appeared first on The Decoder .
- NVIDIA launches Nemotron 3.5 Lightning
NVIDIA unveils its new Nemotron 3.5 Lightning model, a high‑performance LLM aimed at advanced AI applications.
- NVIDIA Nemotron 3.5 Lightning
30B open model for agents, available on Ollama.
- Nvidia's open Nemotron 3.5 Lightning model is all about specialized, local agentic AI
Our AI Model Release Tracker keeps new models in context with their peers, so you know which are worth your time.
- Nvidia releases Nemotron 3.5 Lightning and NeMo Switchyard to give enterprise AI capability options
Artificial intelligence silicon and software giant Nvidia Corp. today announced two new services: a highly customizable Nemotron model and an agentic AI model router named NeMo Switchyard. As enterprises find themselves drowning in artificial intelligence model options, the question is no longer raw power and capability, but fit-for-what-purpose and when. As agents become the norm, […] The post Nvidia releases Nemotron 3.5 Lightning and NeMo Switchyard to give enterprise AI capability options appeared first on SiliconANGLE .
- Nvidia Drives Bang For The Buck With New GenAI Model And Router
Nvidia Drives Bang For The Buck With New GenAI Model And Router
- Meta Launches Muse Glimmer AI Model for Local Coding, Agentic Tasks and Multi-Step Reasoning
Meta has announced Muse Glimmer. This latest AI model from Meta Superintelligence Labs has the model weights under a permissive Apache 2.0 license. It is a 30-billion-parameter model optimised for always-on local agent workflows, and it can work on a Mac or PC with a single consumer GPU. This would enable use cases including local agents and function calling, local co...
- Meta’s local AI model prompts enterprises to rethink hardware-software cost trade-off
Meta rolled out a new 30-billion-parameter AI model optimized to run on a PC or Mac with a single GPU on Monday, offering a way to run always-on agentic workflows locally rather than relying on the cloud. The company has dubbed it Muse Glimmer. However, its hardware demands, including a GPU with a minimum of 24GB of VRAM, could make it difficult to justify for deployment at scale. Although analysts and consultants agree that there is a tremendous enterprise appetite for running models locally, determining whether switching more systems from cloud to local makes fiscal sense is much more complex. The hardware costs are tricky to calculate even today, with the VRAM needed depending on the particular applications to be run. But the far bigger consideration is that there is no way to determine what RAM costs will look like over the next 12-18 months, and there is an identical lack of visibility into how cloud prices might increase during the same timeframe. That makes determining the better financial choice impossible. Agents become capex, not opex Noah Kenney , principal consultant at Digital 520, noted that since RAM costs have increased “exponentially” over the last 12 months, cloud AI providers are also going to have to increase their prices. Beyond that, IT needs to anticipate logistical issues; key questions to ask are, “How quickly can you scale? Can you even get the hardware?” “Meta just made agents a capital expense instead of an operating one,” Kenney said. “For two years, enterprises have been trained to rent intelligence by the token from someone else’s data center. Muse Glimmer runs the agent on a GPU you own, on the desk, with the meter switched off. That is a direct shot at the business model that cloud AI vendors are built on, and it comes from the one player with no cloud API revenue to protect.” In its post announcing the new model, Meta pointed out that it has aggressively slimmed it down to try to make it efficient and cost-effective. “At full precision, a 30-billion parameter model would require over 55 GB of memory — far more than any consumer GPU offers,” Meta said. “We use quantization techniques to compress the model’s weights to approximately 4-bit precision, shrinking the language model to under 20 GB. This leaves enough headroom for the model’s working memory, its KV cache, the perception encoder for image understanding, and the speculative decoding drafter to run simultaneously within a 24 GB or 32 GB envelope. We validated that this compression introduces minimal to no degradation on agentic tasks.” More options for enterprises Mike Wilkes , enterprise CISO at Aikido Security, said that Meta’s move is significant in that it starts to give enterprises more options. “The most important thing about Muse Glimmer is not that Meta has produced another capable model, it is that the economics and architecture of AI are beginning to move back toward the edge,” he said. “The financial comparison therefore becomes capital expenditure that can be amortized over several years versus an effectively perpetual per-token or per-request cloud operating expense.” That means, he said, that an enterprise may rationally pay somewhat more for hardware if doing so gives it predictable AI costs, offline availability, control over model versions, freedom from sudden API pricing or access changes. Independent cybersecurity and risk advisor Steven Eric Fisher also noted that the specs published by Meta don’t tell the full story. The problem is that running locally versus in the cloud can generate a lengthy list of related expenses. “Agentic workloads [in the cloud] can amplify consumption through reasoning, retries, tool calls, context growth, and evaluation, while local deployment [also] introduces hardware, power, lifecycle, support, and utilization costs,” he said, adding that even the RAM requirements need a lot of context. “Meta’s stated 24GB and 32GB memory targets demonstrate that Glimmer can be loaded and executed on comparatively accessible hardware, but that is not the same as having sufficient capacity for meaningful agentic workloads,” Fisher said, pointing out that once other factors are considered, practical memory requirements can move beyond the 32 GB available on an Nvidia RTX 5090 GPU. “In enterprise terms, this still places Glimmer primarily in high-end developer, data science, or dedicated AI workstations rather than the standard corporate desktop or laptop,” he said. Better ROI not guaranteed Justin Greis , CEO of consulting firm Acceligence, agreed. “I wouldn’t assume that moving inference from the cloud to the endpoint automatically produces a lower total cost of ownership,” he said, noting that variable cloud costs would be traded for the price of deployment, endpoint management, support, security, model updates, and potentially accelerated hardware refresh cycles for the local devices. “Muse Glimmer is an important milestone because it makes local agentic AI technically viable. But technical viability and enterprise ROI are two different milestones,” Greis said. “I think Meta has crossed the first one. I do not think they have fully crossed the second one yet.” However, Kenney argued that there are also other elements of the Meta rollout that make meaningful comparisons difficult. “It is worth remembering that the local model is quantized, compressed to roughly 4-bit precision, while the cloud APIs you are comparing against typically serve full-precision models, so this is not a pure apples-to-apples cost comparison,” he said. “The ROI question is not local versus cloud on price alone. It is also a question of whether a company is willing to use a quantized model for the specific task. A cheaper agent that needs more retries or human correction can erase its savings fast.” In addition, a local model often neglects to include every service that a cloud provider typically delivers. “The cloud vendor was quietly handling updates, scaling, reliability, and security patching across your whole footprint. Bring the model in-house and every one of those becomes your problem, multiplied by every machine running it,” Kenney noted. “Most enterprises that consume AI as a service have no muscle for operating a fleet of local models, and that cost rarely gets adequate consideration in ROI conversations. The GPU is cheap. Patching a thousand of them is not.” Sanchit Vir Gogia , chief analyst at Greyhound Research, also stressed that it can be difficult for IT to comprehensively anticipate the different cost variables. “Finance leaders are right to feel skeptical. An engine bought for one employee burns capital whether or not it runs. Meta has shown that a thirty-billion-parameter agent can run on a single machine, and that is a real engineering result. What it has not shown is that such an agent works reliably at enterprise scale, or that a fleet of them can be operated safely. Model fit and production fit are different claims. A laptop must still run the employee’s actual job,” Gogia said. “The economics turn on the incremental hardware premium, refresh timing and actual utilization, measured against the price of the remote inference being displaced.” On the other hand, Arun Chandrasekaran , a distinguished VP analyst with Gartner, said that he found it “very interesting that they have decided to release a smaller model that operates on the edge” and especially liked Meta’s use of the popular Apache license. But he would have preferred that they had released more than just a model. “Enterprise customers are asking for a car and Meta is delivering an engine,” Chandrasekaran said. “They should have built something more like a platform solution.”
- Meta Superintelligence Labs unveils on-device model Muse Glimmer
Meta has open sourced Muse Glimmer, a 30 billion-parameter model under Apache 2.0 that runs locally on a single consumer GPU.
- Meta launches open source AI model called Muse Glimmer for devices
Meta launches open source AI model called Muse Glimmer for devices warriorswire.usatoday.com
- Meta launches open source AI model called Muse Glimmer for devices
Meta launches open source AI model called Muse Glimmer for devices Chiefs Wire
- Meta unveils an open version of its most powerful AI model
Meta unveils an open version of its most powerful AI model The Boston Globe
- Zuckerberg manifesto pushes an open-source approach on AI as Meta releases its latest model
Zuckerberg manifesto pushes an open-source approach on AI as Meta releases its latest model Inquirer.com
- Zuckerberg manifesto pushes open-source approach on AI as Meta releases latest model
Meta Platforms has released a new AI model for developers, emphasizing open-source access
- Meta’s new AI model runs entirely offline, but your GPU needs to keep up
Meta released Muse Glimmer, a free 30B parameter AI model you can run entirely offline. No subscription, no data center, just a GPU with at least 24GB of VRAM.
- OpenAI Launches GPT-5.6-Cyber Model Through Daybreak to Strengthen Cyber Defence
OpenAI announced the GPT‑5.6‑Cyber model, available through Daybreak Red access. It is designed to handle advanced cybersecurity tasks with fewer refusals. The company states that it has conducted an internal test and outperformed its general-purpose models. The company offered GPT-5.6-Cyber access to trusted customers, who used the model to improve their defensiv...
- OpenAI launches GPT-5.6-Cyber to find vulnerabilities before hackers
OpenAI launches GPT-5.6-Cyber to find vulnerabilities before hackers YourStory.com
- OpenAI expands Daybreak with GPT-5.6-Cyber model as AI cyber breaches surge
OpenAI's GPT-5.6-Cyber targets advanced cybersecurity tasks, including vulnerability research, exploit validation and security testing
- OpenAI launches GPT-5.6-Cyber, expands Daybreak initiative with 2 new tiers
OpenAI launches GPT-5.6-Cyber, expands Daybreak initiative with 2 new tiers
- OpenAI extends 'Daybreak' security project and reveals new cyber model — but for approved users only
Daybreak now offers two different models that come with varying degrees of compliance.
- Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning
Learning reliable surgical manipulation policies is bottlenecked by the scarcity of action-labeled demonstrations: teleoperated surgical robot (e.g., dVRK) trajectories with synchronized kinematics are costly to collect, while surgical tasks demand precise contact handling, long-horizon reasoning, a...
- Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation
GUI Visual Grounding is a fundamental capability for GUI agents. Existing models typically freeze their parameters after deployment, limiting their ability to adapt to unseen interfaces. Although recent methods attempt to adapt models via test-time reinforcement learning, they cannot reflect upon fa...
- How to Verify Consistency of Probabilistic Claims
When a probabilistic predictor answers many conditional-probability queries, are its answers self-consistent, and can this be verified in polynomial time? This problem is of interest for AI safety, where safety is derived from honesty about probabilistic predictions of unwanted outcomes potentially ...
- From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop
The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc interpretability of static models to mechanistic understanding and proa...
- Entropy-Centric Explainable AI for Remote Sensing Image Segmentation
Artificial intelligence (AI) has become a powerful approach to solving complex problems in critical domains. Many concerns arise regarding the decision-making process of its models, mainly due to deep neural networks outperforming their peers at the cost of ambiguity in feature extraction and predic...
- A Comparative Evaluation of Deep Learning Object Detection Models on a Real-World Multi-Plant Dataset from Africa
The application of computer vision in agriculture has shown significant potential for improving crop monitoring and precision farming. However, many existing approaches rely on controlled datasets that do not adequately represent realworld farming conditions, particularly in underrepresented regions...
- 3D Weighted Geometric Graph Neural Networks for Sheep Facial Pain Assessment
Deep learning systems perform mainly within the 2D for a single image domain and take the face as a single-dimension representation, losing sight of the 3D anatomy of sheep and cross-landmark spatial relationships that are intrinsic to the clinically proven Sheep Pain Facial Expression Scale (SPFES)...
- Multiclass Sentiment Analysis for Identifying Political Viewpoints
The rapid growth of social media has created vast amounts of political discourse, which provides valuable opportunities to analyze public opinions and identify different political perspectives. Sentiment Analysis (SA) is a core task in Natural Language Processing (NLP) that allows the computational ...
- V-FiLLM: Verified Financial LLM Reasoning Benchmark
While existing benchmarks have made substantial progress in evaluating LLMs across STEM domains, financial reasoning over structured data remains comparatively less explored. We introduce V-FiLLM, a framework that generates financial reasoning benchmarks from executable computation trees grounded in...
- Workflow Cards: Structured Summaries of Workflow Executions Using Provenance Data
Model Cards and Data Cards have demonstrated the value of structured, human-readable documentation for machine learning artifacts, capturing their context, parameters, limitations, and intended use. However, these practices remain focused on static artifacts (the datasets and trained models themselv...
- R4DSG: Relative 4D Scene Graph Memory for Object-Centric Question Answering in Long Egocentric Video
Long-horizon egocentric video is a rich substrate for wearable AI assistants, but object-centric questions such as where an item was moved, when it last changed state, or why it was relocated remain difficult because caption- and transcript-based memories rarely preserve persistent object identity o...
- Putting Registers to Work: Task Registers for Token Pruning in Vision Transformers
Token-pruning policies are usually designed for a single recognition pipeline, but pretrained Vision Transformers are reused across tasks with different spatial demands. We ask which parts of a pruning policy transfer across image classification, semantic segmentation, and object detection. For each...
- TimeRoute: Time-Aware Modality Routing and Diffusion for Multi-Modal Recommendation
Multi-modal recommenders fuse collaborative signals with item modalities such as text, images, and audio, but the usefulness of each drifts over time and at different rates. For example, chocolate purchases typically guided by textual ingredient cues can shift toward visual packaging and ambient aud...
- XCoT-VLA: Executable Chain-of-Thought for Vision-Language-Action Driving
Vision-Language-Action (VLA) models can connect scene understanding, semantic reasoning, and trajectory generation for autonomous driving. However, verbose natural-language Chain-of-Thought (CoT) is poorly suited to real-time control because it is open-ended, costly to decode, and difficult to optim...
- ReLTEx: Reliable LLM-based Taxonomy Expansion
Recent advances in Large Language Models (LLMs) have demonstrated strong capabilities in generating semantically relevant concepts and relations, making them promising tools for taxonomy enrichment. However, directly relying on LLM-generated expansions often leads to noisy, redundant, or hierarchica...
- CARE: Confidence-Aware Reasoning for Reliable Medical VQA
Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual question answering, yet these models suffer from $\textit{confidence miscalibration}$---a systematic gap between expressed certainty and actual diagnost...
- Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes
While Multimodal Large Language Models (MLLMs) demonstrate impressive performance in benign scenarios, their cognitive reliability deteriorates significantly in complex scenes under adverse conditions. In these settings, models often rely on implicit inference without sufficient visual evidence, lea...
- A Cost-Efficient Routing Pipeline for Multilingual Short-Text Classification Using Small Language Models
Multilingual short-text classification supports operational systems such as content moderation, customer support routing, and intent recognition, yet aggregate evaluation often hides large differences between high-resource and low-resource languages. Uniform inference policies are simple to deploy, ...
- Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation
Understanding camera motion is fundamental to video perception, with applications in spatial intelligence and controllable video generation. Multimodal large language models (MLLMs) provide a natural interface for this task, but existing work typically assigns one or more labels to an entire clip. S...
- FaithformBench: Benchmarking Faithfulness of Mathematical Chain-of-Thought Autoformalisation
Autoformalisation (AF) systems map natural language reasoning steps into formal statements in a proof assistant such as Lean. We consider how to assess the faithfulness of these systems. Existing approaches require expensive human-annotated ground truth, or rely on LLM judges or embedding models, wh...
- VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?
Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments. Everyday life assistance is different. A task runs for weeks rather than minutes. The world keeps changing while the ag...
- Hypothesis Frontier: Verifier Guided LLM and Symbolic Search for First-Order Induction
First-order concept synthesis asks a system to infer one formula that classifies labeled objects consistently across several finite relational structures. Every candidate can be evaluated exactly, but quantified first-order formulas form a vast search space, and LLM outputs are often semantically pr...
- Whisper-Aware LLM: Self-Supervised Uncertainty Learning for Robust Whispered Speech Recognition
The signal ambiguity of whispered speech drives ASR systems toward two opposing failure modes: failing to capture whispered speech or hallucinatory transcription of noise. This paper introduces the Whisper-Aware LLM, a framework that teaches an Audio-LLM to perceive and react to this uncertainty. Ou...
- BPG: Balancing Plasticity and Generalization for Domain Incremental Learning
Deep neural networks excel in various tasks but struggle to generalize across evolving data distributions, leading to significant performance degradation under domain shifts. Domain incremental learning (DIL) addresses this challenge by enabling models to continuously adapt while retaining prior kno...
- EvoMem: Memory-Augmented Evolution for Code Optimization
Successful mutation strategies in evolutionary code search may contain reusable knowledge that is useful beyond a single run, and in some cases may transfer across related tasks and domains. However, existing LLM-driven evolutionary frameworks largely discard such knowledge, repeatedly rediscovering...