AI News Archive: July 8, 2026 — Part 14
Sourced from 500+ daily AI sources, scored by relevance.
- OpenAI To Launch New Model After US Freeze
OpenAI To Launch New Model After US Freeze Barron's
- ChatGPT's New Voice Models Can 'Listen' and 'Talk' at the Same Time
OpenAI says the new AI models should be better at live translation.
- OpenAI's GPT-5.6 Is Dropping on Thursday: What's Different About Sol, Terra and Luna
OpenAI says its flagship model for ChatGPT should make fewer mistakes.
- OpenAI’s New GPT-5.6 Is Coming Sooner Than You Think
OpenAI is set to launch GPT-5.6 after a US security review, raising new questions for enterprise AI access, safeguards, and governance. The post OpenAI’s New GPT-5.6 Is Coming Sooner Than You Think appeared first on TechRepublic .
- You’ll finally be able to try OpenAI’s GPT-5.6 Sol, Terra, and Luna models this week
After nearly two weeks of limited preview access, OpenAI is finally ready to roll out GPT-5.6 Sol, Terra, and Luna to the public on July 9.
- agents-cli
The CLI your coding agent uses to ship agents
- OpenAI shares update on GPT-5.6 availability after holding back release
OpenAI unveiled GPT-5.6 at the end of June, but the public release was held back while the U.S. government reviewed the new AI models. Now the company has shared an update on when customers will be able to use the upgraded model: this week.
- OpenAI Introduces GPT-Live to Make ChatGPT Voice Feel Like a Real Conversation
OpenAI today introduced GPT-Live, which it describes as a new generation of voice models meant to make talking to AI feel more like having a conversation with a real person. GPT-Live is meant to replace the existing ChatGPT voice experience. GPT-Live is able to listen and speak at the same time, and it can show it is paying attention with acknowledgment phrases like "mhmm." The model was built for continuous interaction, and it can make decisions on whether to speak, continue listening, pause, interrupt, or use a tool multiple times per second. Talking with ChatGPT should now feel much more like a real conversation. You can interrupt with a question, pause to gather your thoughts, or ask ChatGPT to slow down. It naturally acknowledges what you're saying with phrases like "mhmm" or "got it," so you know it's following along. We've also remastered the nine distinct voices in ChatGPT for GPT-Live. OpenAI says GPT-Live is its smartest voice model to date, using the latest frontier model (currently GPT–5.5) for web search, deep reasoning, and complex work. While GPT-Live works on a task, it is able to continue a conversation, and then give the results of a task when it's finished. It also works for live translation, and displays rich visual cards for weather, stocks, sports, and more. OpenAI is rolling out GPT-Live–1 and GPT-Live–1 mini to ChatGPT users worldwide starting today. GPT-Live–1 is the default for Go, Plus, and Pro users, while GPT-Live–1 mini is the default for Free users. ChatGPT users can tap the Voice button to talk with ChatGPT and experience GPT-Live. GPT-Live does not yet support voice with video or screen sharing in ChatGPT, but OpenAI is working to add that feature soon. Tags: ChatGPT , OpenAI This article, " OpenAI Introduces GPT-Live to Make ChatGPT Voice Feel Like a Real Conversation " first appeared on MacRumors.com Discuss this article in our forums
- OpenAI to launch new model after US freeze
ChatGPT-maker OpenAI said its latest and most powerful artificial intelligence model will be released to the public on Thursday, as the U.S. government reportedly approved a broader launch.
- OpenAI’s advanced GPT-5.6 models to be publicly released
After working alongside government partners for safety evaluations, OpenAI said it is “expanding preview access globally” for its latest powerful model series.
- Exclusive: China’s MiniMax Plans to Launch 2.7-Trillion Parameter Model
Exclusive: China’s MiniMax Plans to Launch 2.7-Trillion Parameter Model The Information
- China's MiniMax plans to launch giant 2.7 trillion parameter model
China's MiniMax plans to launch giant 2.7 trillion parameter model Reuters
- China's MiniMax plans to launch giant 2.7 trillion parameter model
Chinese AI firm MiniMax is developing a 2.7 trillion-parameter model. This new model could become the world's largest open-weight AI system. Cheaper Chinese AI models are gaining traction as alternatives to US systems. MiniMax will also launch a multimodal video generation model soon. The company is planning a second listing on Shanghai's STAR Market.
- Meta launches Muse Image to take on OpenAI and Google AI
Meta launches Muse Image to take on OpenAI and Google AI YourStory.com
- Meta’s Muse Image Might be Just What SMBs Need
Meta’s new AI imaging model enables users to create competitive ad content on the social media giant’s platforms.
- Meta's new image model outranks competitors
Meta releases a new image generation model that surpasses current leading competitors in performance.
- Meta launches Muse AI image
Meta unveils Muse AI, a new image generation tool.
- Meta's New Image Model Lets People Tag Your Insta ID in Prompts: How to Opt Out
Meta's New Image Model Lets People Tag Your Insta ID in Prompts: How to Opt Out PCMag UK
- Meta launches image generation model with coding, search capabilities
Meta Platforms Inc. today debuted an image generation model that can write code and search the web. Muse Image is the second algorithm released to date by Meta Superintelligence Labs, the company’s artificial intelligence research group. The first is the Muse Spark large language model that made its debut in April. Both algorithms are available […] The post Meta launches image generation model with coding, search capabilities appeared first on SiliconANGLE .
- Meta rolls out Muse Image generative AI model
Meta Platforms has launched Muse Image, its image-generation model developed by Meta Superintelligence Labs.
- Meta rolls out Muse Image generative AI model
Meta rolls out Muse Image generative AI model verdict.co.uk
- Meta’s Muse Image Explained: What to Know About the New AI Image Model
Meta Muse Image brings AI image generation to Instagram, WhatsApp, and Meta AI. Here’s what users should know about features and privacy. The post Meta’s Muse Image Explained: What to Know About the New AI Image Model appeared first on TechRepublic .
- Ant Group’s Robbyant open-sources robotics AI models
Robbyant said LingBot-VLA 2.0 is undergoing commercial pilot testing in retail sorting, logistics, and industrial automation.
- Co-LMLM: Continuous-Query Limited Memory Language Models
Limited memory language models (LMLMs) externalize factual knowledge during pretraining to a knowledge base (KB), rather than memorizing it in their weights. During generation, the model then fetches knowledge from the KB as needed. This recently introduced paradigm provides multiple advantages, inc...
- Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety
We introduce institutional red-teaming, an evaluation methodology for testing deployment rules in multi-agent AI: hold the agents, objectives, and task state fixed, vary only one rule, and attribute the resulting change in collective behavior to that rule. We instantiate the methodology in IABench-C...
- Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF
Reinforcement learning from human feedback (RLHF) has emerged as a powerful paradigm for aligning generative models with human preferences. However, applying RLHF to diffusion models remains highly feedback inefficient, as existing approaches typically require large amounts of human or reward model ...
- Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning
Reinforcement learning from verifiable rewards (e.g. GRPO) is the engine behind today's reasoning models, yet it grades only the final answer. On hard problems this trains models to write more rather than to think better, since the trace itself is never graded and no label for good thinking exists. ...
- RL Post-Training Builds Compositional Reasoning Strategies
Does RL post-training merely amplify primitive skills already latent in a base model, or can it compose primitive skills into new higher-level strategies? We study this question in a fully observable rewrite-grammar environment where the pretraining distribution is known and every generated rewrite ...
- ALER-TI: Aligned Latent Embedding Retrieval for Time Series Imputation
Deep learning has significantly advanced time series imputation, yet most existing architectures primarily rely on localized temporal context within the corrupted input sequence. This reliance can be limiting in real-world scenarios, where time series often exhibit non-stationary dynamics, weak temp...
- Future Confidence Distillation in Large Language Models
Reliable confidence estimation is essential for deploying large language models (LLMs) in confidence-aware systems, where downstream decisions such as retrieval, tool use, and adaptive computation depend on accurately estimating answer reliability. Existing approaches, however, largely treat confide...
- Towards Agentic AI Governance: A Preliminary Assessment
Artificial intelligence is rapidly evolving from generative systems to agentic AI capable of autonomously planning and executing tasks. Widely characterized as the Year of Agentic AI, 2025 marked accelerated development and deployment, introducing new ethical and governance challenges. This paper pr...
- CARLA-GS: Decoupling Representation, Reasoning, and Physics Simulation for Autonomous Driving Corner-Case Synthesis
Safety evaluation for autonomous driving is dominated by rare, safety-critical interactions, motivating simulators that can deliberately synthesize corner cases with photorealistic observations. Corner-case generation is inherently a multi-source problem spanning visual representation, scene reasoni...
- Collaborative Synthetic Data Generation for Knowledge Transfer in Federated Learning
One-shot federated learning (OSFL) addresses the communication overhead of federated learning by limiting training to a single round, but doing so without sacrificing model quality is non-trivial, particularly when client data distributions diverge. Recent work has addressed this challenge by aggreg...
- Stability of Flow Models for Graph Signals
Generating signals on graphs requires permutation-equivariant models that exhibit stability with respect to relative structural perturbations. While favorable stability properties of Graph Neural Networks (GNNs) have been well documented, it is unclear how structural errors propagate through the dyn...
- Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
Reinforcement learning (RL) is becoming increasingly important for post-training large language models (LLMs). Previous RL pipelines for LLMs were mostly synchronous and batch-interleaved, which is inefficient for long-horizon agentic tasks. Recently, asynchronous RL has emerged as a more efficient ...
- HIVE: Understanding Post-Hallucination Reasoning in Vision Language Models
Hallucinations in vision language models (VLMs) are commonly treated as semantic errors, yet they often arise from partial or ambiguous visual evidence. Prior work mainly focuses on detecting or suppressing hallucinations at generation time, leaving the subsequent reasoning stage largely unexplored....
- Do LLM-Generated Skills Make Better AI Data Scientists? A Component Ablation Across Data-Science Workflows
Product data scientists often ask LLM-based agents to help with recurring execution tasks such as cleaning data, writing SQL, choosing statistical tests, and formatting results. Reusable skill files are meant to avoid prompting from scratch by packaging guidance for a task family. Expert-written ski...
- TimEE: End-to-end Time Series Classification via In-Context Learning
Time series classification (TSC) is dominated by a two-stage paradigm: train a feature encoder -- either from scratch on the target dataset or via pretraining on large corpora -- and then fit a task-specific classifier on top. While effective, this decoupling optimizes representation learning indepe...
- Where to Intervene? Benchmarking Fairness-Aware Learning on Differentially Private Synthetic Tabular Data
Machine learning models are increasingly deployed in high-stakes domains, raising concerns about both privacy and fairness. Differential Privacy (DP) has become a gold standard for privacy-preserving data analysis, while fairness-aware mechanisms aim to mitigate discrimination against underrepresent...
- SynthAVE: Scalable Synthetic Labeling for E-Commerce with LLM-Arena Validation
Fine-tuning large language models (LLMs) for e-commerce attribute extraction requires labeled data representative across thousands of product types, attributes, and multiple languages. This combinatorial scale translates to millions of annotations, rendering human labeling prohibitively costly. Whil...
- SpaCellAgent: A Self-Evolving LLM-Based Multi-Agent Framework for Trajectory Analysis
Spatial and Single-cell transcriptomics are transformative in deciphering cellular dynamics. As the fundamental paradigm for reconstructing cell developmental paths, trajectory inference (TI) is critical. However, existing methods require extensive manual intervention and proficiency in heterogeneou...
- Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents
Tool-using LLM agents can violate the very policies they are deployed to enforce while appearing to complete the task successfully. In policy-permissive environments, a tool may execute any well-formed call even when the corresponding state transition is forbidden by domain policy. The result is a s...
- Heterogeneity-Adaptive Diffusion Schrodinger Bridge for PET-Guided Whole-Body MRI Translation
While whole-body multimodal medical imaging scanners have been increasingly recognized for more effective medical applications, the excessive long acquisition time in PET-MR scanning is a major obstacle in more efficient clinical practice. Deep learning-based MRI translation provides a potential sol...
- When Prompts Ignore Structure: Graph-Based Attribute Reasoning for Calibrated VLMs
Reliable confidence estimation remains a key limitation of test-time adaptation in vision-language models (VLMs), where prompt tuning improves zero-shot accuracy but often degrades calibration due to entropy-driven overconfidence. Prior approaches mitigate this using LLM-derived class attributes and...
- Physics-Audited Agentic Discovery in Scientific Machine Learning
In agentic scientific machine learning (SciML), large language model (LLM) agents can discover surrogate models and select one by an automated score, typically an error metric. A low error, however, does not establish that the predicted fields satisfy the physics that matter for mechanics, such as b...
- On Adversarial Vulnerability of Vision-Language Models through the Lens of Intermediate Spectral Subspaces
Adversarial vulnerability in deep neural networks (DNNs) has been studied from the perspectives of decision-boundary geometry, feature robustness, input-output Jacobians, and the instability of inverse problems. Here, we focus on the spectral structure of intermediate linear transformations that pro...
- Hypergraph Neural Stochastic Diffusion: An SDE Framework for Uncertainty Estimation
Hypergraph neural networks have shown powerful capability in modeling higher-order relations, yet their predictive uncertainty remains underexplored. Unlike pairwise graphs, uncertainty in hypergraphs arises not only from noisy attributes and ambiguous labels, but also from variations in node-hypere...
- From Atomic Actions to Standard Operating Procedures: Iterative Tool Optimization for Self-Evolving LLM Agents
Tool utilization enables Large Language Model (LLM) agents to interact with the real world and resolve complex tasks. However, existing agent frameworks predominantly rely on static toolsets composed of granular atomic actions (e.g., basic file I/O or single-turn search), which forces agents to rein...
- FedCVESA: Taking Away Training Data in Federated Learning via Correlation Value Encoding and Segmented Aggregation
Federated learning (FL) avoids explicit data exposure by keeping raw data on local clients, yet privacy risks remain in the training process and the learned model itself. Recently, centralized Taking Away Training Data (TATD) attacks have shown that malicious training could abuse the memorization ca...
- Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders
Turn-taking prediction is a key requirement for social robots involved in human-human interaction, particularly in mediator settings, where the robot must anticipate conversational dynamics rather than merely react to pauses. This work presents a Multimodal Voice Activity Projection (MM-VAP) framewo...