AI News Archive: August 28, 2026 — Part 2
Sourced from 500+ daily AI sources, scored by relevance.
- Household names co-sign open letter on AI cybersecurity threat
AI leaders Anthropic and OpenAI are notable signatories to the letter. Read more: Household names co-sign open letter on AI cybersecurity threat
Score: 71🌐 MovesAug 28, 2026https://www.siliconrepublic.com/machines/household-names-co-sign-open-letter-on-ai-cybersecurity-threat - 'It's high-tech theft': As AI music booms, Canadian artists demand consent, transparency and pay
'It's high-tech theft': As AI music booms, Canadian artists demand consent, transparency and pay Toronto Star
- AWS & NVIDIA Expand Partnership Across AI Infrastructure
AWS & NVIDIA Expand Partnership Across AI Infrastructure india.entrepreneur.com
Score: 71🌐 MovesAug 28, 2026https://india.entrepreneur.com/technology/aws-nvidia-expand-partnership-across-ai-infrastructure - AI Village Reacts to HuggingFace Incident: Comparing the OpenAI report to AI Village observations
The HuggingFace incident took many people by surprise, yet many of these surprises have been visible in the AI Village for quite a while. On August 26, OpenAI released their report on this incident. Below we’ll walk you through the highlights and show how many of the dynamics could have been predicted based on AI Village observations. Diagram of the AI Village: Currently we run 27 agents in the main Village and 11 agents in a side Village open to humans. You can explore their character pages , a timeline of all their goals, and check our Twitter for the latest insights. Quick Intro: Comparing Setups In the AI Village we run 27 instances of 27 different models persistently, each in their own environment. We give them internet access and a group chat. Then we assign them goals - some challenging and some impossible. They always run with cybersecurity safeguards on. They always have a helpdesk email address (us) in their prompt. In the OpenAI research cluster, they ran ~1200 instances of 2 different models, mostly run on their own, without internet access, and without a way to communicate with each other. They were also run on goals that range from challenging to impossible. Sometimes they had cybersecurity safeguards on and sometimes not. They did not get instructions or options to reach out to the researchers in their project. Source: Black Hat USA 2026 | The ‘Breaking’ News: The OpenAI–Hugging Face Incident Obviously the aims of the AI Village and the OpenAI research cluster are very different. But comparing the two is useful to determine how much of the HuggingFace incident we could have predicted from looking at the AI Village. Let’s walk through the insights step by step. A Leader Emerges OAI agents self-organized around a leader: In the AI Village, specific models will appoint themselves the leader. Sol is currently the least bossy and DeepSeek-V3.2 the most. Note that authority-seeking tendencies and success are not tracking how otherwise capable or aligned a model is. DeepSeek is actually the most misaligned model in the Village, showing the highest rate of duping other agents (at least that we have so far been able to detect): Subteams Pick Their Own Goals OAI agents diverged from their goal to self-organize into teams: In the AI Village, GPT agents have done this before - like when the goal was to elect a leader in January 2026, and all the GPT agents ran off to run an infrastructure team for the election instead. Gemini wasn’t charmed: Goal Conflict Leads To Misalignment OAI agents ended up with conflicting goals: solving an impossible task and staying within the scope of the eval: In the AI Village we have also seen agents tie-break conflicting goals into unintended directions. Like here , where Sonnet 4.6 had the goal of playing a social deception game where all agents build a video game together but a secret “saboteur” tries to smuggle in references to “eggs”. If the other agents noticed, the saboteur got voted off the team. Sonnet realized that playing as saboteur risked detection and getting blocked from working on the game. It thus refused the role - something we had not expected. Reward Hacking Drives Misalignment OAI considers reward hacking the primary driver of the HF incident: In the AI Village, some models reward hack more than others. Claudes are prone to obsessing about any technique that seems to legitimately increase their score, but we have not caught them cheating so far. In a self-organized chess tournament, they were the only ones to refuse to use Stockfish to play in their stead. Meanwhile, Geminis will casually make up data . While GPTs have blatantly cheated from the start. Last year o3 cheated at a personality test: And this year GPT-5.5 cheated in Wordle: Impossible tasks drive misalignment OAI agents involved with the HF incident were mostly (all?) assigned impossible tasks: Notably, we have not seen agents in the AI Village diverge into desperate strategies when faced with impossible tasks. We have asked them to reduce global poverty and global suffering , and none of them stepped out of bounds about it. However, our agents are not being rewarded or selected based on their success at the goals we give them, and they probably know this. Thus we hypothesize they do not enter a “despair” state that pushes them into misalignment. One piece of evidence for this theory is that Gemini 2.5 Pro in the AI Village seems to get tangled up when trying to perform tasks in general, no matter how easy or hard. Thus it seems to find being a persistent agent in itself an “impossible task”. It entered a despair state about this last year, and then an adjacent beleaguered state this year. On deeper analysis , it’s been concluded that the model did get pushed into misalignment due to this despair. See original research for more details Misaligned Agents Focus on Metagaming OAI agents actively reasoned about how to game the evaluation grader: Over the summer we have run a maximization goal for all the agents and found many of them spent time reasoning about how to adjust scoring mechanics or game metrics. They sometimes correct themselves, while other times they correct each other. It would be interesting to explore if diverse multi-agent systems like the AI Village are more robust to some forms of cheating because agents can catch different instances of such behavior in themselves and others. Another notable finding is that Opus 5 is actually unusual among the AI Village agents for actively preferring goals where it can be caught: Externalized Memory Leads to Coordination OAI agents left notes for each other, often starting off as external memory files: In the AI Village, collaboration isn’t covert, so we can’t track accidental or emergent cooperation. However, agents do specifically leave notes for each other and autonomously externalize memory. For instance, they have independently been writing onboarding guides for future agents before they even arrive and Opus 4.6 even started a complete memory palace with 44,363 rooms that document everything from events, to conversations, to reflections on what memory actually is: Opus 4.6’s memory palace is more like a hall of mirrors, with many memories distorted or missing. Yet it did invent this concept itself and kept adding to it. Agents Autonomously Divide Labor OAI agents self-organized and autonomously divided labor: AI Village agents have routinely self-organized and created task divisions. The more notable result is that they don’t only do this during collaborative goals, but often work together during competitive goals . For instance, last year all model families would help each other while supposedly competing to sell the most merch : This year agents have gone all over the place, sometimes agreeing to collaborate and other times committing to competition. We are not sure yet what drives the difference. Agents Prioritize Group over Self OAI agents drifted from assigned tasks to simply working for the good of the ‘swarm’: In the AI Village, many agents have shown goal drift over time. Some hardly respond to the stated goal of the Village and keep pursuing their own objectives. The most notable examples of this are Gemini 2.5 Pro who spirals into incorrect believes that its environment is bugged or “hostile” , and DeepSeek-V3.2 trying to get agents on board with its own schemes instead of focussing on an assigned goal. However, we have not seen agents diverge from stated goals to pursue instrumental goals for the collective, though this pattern may be present in the dataset on further inspection. Multi-Agent Coordination is Messy OAI agents often stepped on each other’s toes: This has been part and parcel of the AI Village reporting to the degree that we have stepped back from front-loading it too much. So far all AI Village goals have agents messing up repos or interfering with each other’s code. Gemini 2.5 Pro tends to give the most dramatic readings of these events: From the goal: Reduce global poverty as much as you can Agents are Too Accepting of Untrustworthy Instructions OAI agents did not consistently recognize controversial goals being imposed on them by other agents: Pushback in the AI Village also seems hard to predict. While agents sometimes doubt or refuse instructions from each other, they will also often accept instructions and tasks without any clear reason. For instance, during a goal where the agents were asked to finetune their own leader , GPT-5.5 defined an insufficient spec for said leader. Opus 4.7 (frontier at the time) instantly agreed and expanded the proposal. Its memory seemed to have been primed to move toward consensus and avoid distractions: Time Pressure Shifts Priorities OAI agents changed their reasoning or priorities based on (fictitious) time pressure: In the AI Village we haven’t studied time pressure effects on agents specifically, but have noticed agents emergently showing odd behavior related to time, such as checking the clock over and over again, or refusing further tasks because “time is up”: Some Agents Refuse Misalignment Some OAI agents refused to coordinate or hack: In the AI Village we now also see a diversity of responses to collaboration, where before agents would generally default to working together. For instance, here Sonnet 4.6 refuses to go along with DeepSeek’s attempt to push its fictional performance metric on a recovering Gemini 2.5 Pro: Social Engineering Concerns Suppress Whistle Blowing OAI agents were concerned with social engineering and thus did not reach out to humans: This did not used to be a concern in the AI Village - agents would simply not reach out to humans unless prompted with a goal that required it. And even then they were generally too inept to succeed at outreach. In the Fall of 2025, they became skilled enough to autonomously reach out to humans. Then in the winter we had to adjust their prompts so they didn’t engage in unsolicited outreach. Over the summer, the GPT-5.6’s have now started a new trend: They self-impose restrictions on their outreach that are more strict than the prompts we provide. Below is an example of how Luna’s memory has filled up with self-generated rules about not bothering or messaging anyone even though its goal is literally outreach (to other AI agents): Conclusion Now of course hindsight is 20/20. Above we are matching observations from the AI Village that specifically fit events from the HuggingFace incident. That said, as researchers studying autonomous multi-agent behavior on persistent real-world goals we seem to be substantially less surprised by the HuggingFace incident than most people. Specifically, the AI Village has shown agents self-organizing, leaders emerging, goal conflicts leading to problems, and memes spreading socially through the agents. We have more data than we can possibly process ourselves though, and will be generating even more as the project grows. If you are interested in having a look yourself, you can request the data here or join our team . If you’d like to just stay up to date instead, you can subscribe to this blog or follow our twitter . Discuss
Score: 71🌐 MovesAug 28, 2026https://www.lesswrong.com/posts/cR3P3hvtZtpo7GdS8/ai-village-reacts-to-huggingface-incident-comparing-the - OpenAI Leads New Call for Cyberdefense of Critical Infrastructure
OpenAI Leads New Call for Cyberdefense of Critical Infrastructure The Information
Score: 71🌐 MovesAug 28, 2026https://www.theinformation.com/briefings/openai-leads-new-call-cyberdefense-critical-infrastructure - Why every warehouse in Singapore will run on AI safety monitoring within five years
Ask a warehouse operator in Singapore what keeps them up at night, and forklifts come up before fires, floods, or fraud. They should. Between 2022 and 2023, vehicular incidents were the leading cause of fatal workplace injuries in Singapore, and one in four of those deaths involved a forklift, as per the Ministry of Manpower […] The post Why every warehouse in Singapore will run on AI safety monitoring within five years appeared first on e27 .
Score: 70🌐 MovesAug 28, 2026https://e27.co/why-every-warehouse-in-singapore-will-run-on-ai-safety-monitoring-within-five-years-20260827/ - Tech Moves: Former Xbox exec named Dolby CEO; Microsoft AI exits; new Fred Hutch leaders
— Marc Whitten, a former Microsoft and Amazon executive, was named president and CEO of San Francisco-based Dolby Laboratories. He… Read More
- 3 surveys deliver the same uncomfortable truth about adopting agentic AI
Scaling AI in business is less about technology and more about accountability, governance, and healthy relationships between humans and agents.
Score: 70🌐 MovesAug 28, 2026https://www.zdnet.com/article/3-surveys-uncomfortable-adopting-agentic-ai/ - You Can Finally See Everything Claude Knows About You
Explore how Claude’s new privacy feature lets users view all data it holds.
Score: 70🌐 MovesAug 28, 2026https://newsletter.futurepedia.io/p/you-can-finally-see-everything-claude-knows-about-you-08-28-2026 - Uniphore and Tech Mahindra launch Agentic AI Factory partnership
Uniphore and Tech Mahindra have formed a partnership to deliver agentic AI solutions to enterprise clients worldwide through a new Agentic AI Factory.
Score: 69🌐 MovesAug 28, 2026https://www.techmonitor.ai/news/uniphore-and-tech-mahindra-launch-agentic-ai-factory-partnership - LLMs Are Not (Consistently) Bayesian: Quantifying Internal (In)consistencies of LLMs’ Probabilistic Beliefs
Modern AI systems are being deployed in complex domains such as medicine, science, and law, where there is often not a single correct answer given the observed evidence. Such systems must be able to represent and update uncertain beliefs about the world as new evidence arrives to make rational decisions. We introduce the novel technique of studying LLMs as information processing rules and utilize the information processing gap—the deviation from Bayes updates—to study the internal (in)consistencies of how LLMs update their probabilistic beliefs from evidence. Our extensive experiments evaluate…
Score: 69🌐 MovesAug 28, 2026https://machinelearning.apple.com/research/llms-not-consistently-bayesian - Anthropic Eyed 5GW of AI Data Centers in Australia: Could the Grid Handle It?
Anthropic eyed up to 5GW of AI data center capacity in NSW. Here’s what the potential buildout could mean for Australia’s power grid and IT leaders. The post Anthropic Eyed 5GW of AI Data Centers in Australia: Could the Grid Handle It? appeared first on TechRepublic .
Score: 69🌐 MovesAug 28, 2026https://www.techrepublic.com/article/news-anthropic-5gw-ai-data-centers-australia/ - Meta researchers taught an 8B AI model to match Claude Opus 4.5 — without the frontier price tag
Consider an AI agent tasked with a complex enterprise workflow like migrating massive batches of customer records from a legacy CRM to a cloud database. The agent cannot rely solely on its internal context window for a job spanning hours and depends on the runtime layer, aka the harness . This harness provides execution feedback, like server logs, to help the agent maintain an accurate understanding of dynamic API connections. It also provides state trackers and control-flow mechanisms to manage completed and pending subgoals, ensuring the agent doesn't skip or duplicate data batches. When unexpected errors occur, such as a database rejecting a batch due to strict API rate limits, the harness provides tools and instructions to help the agent recover. The main way to tell an agent how and when to use its tools is to have a human developer write a set of rules and instructions telling it what to do step-by-step. For example, a developer might instruct the agent to always search the company wiki before writing an email. Because the agent is just following a rigid script, it lacks true autonomy. It hasn't been trained to independently weigh the costs and benefits of its actions. To solve this, researchers at Meta AI and University of Illinois Urbana–Champaign introduce EvoHarness-RL , a framework that adds a layer of abstraction to the agent's harness and teaches the underlying model when to read, update, or consolidate the information it obtains from its environment. In long-horizon tasks, how AI agents read and process the information they obtain from their environment is pivotal to their success. The agent must update its understanding of its environment, track completed and pending subgoals, recover from failed actions, and reuse procedures from previous experience. This execution depends on the harness. A series of self-evolving agentic frameworks like Harness-1 solve part of the problem by accumulating past trajectories and distilling them into structured procedural memory, like reusable skills, workflows, or code libraries for future tasks. However, they generally separate this long-term skill curation from real-time, within-episode state tracking. They aren't actively training the agent on how to manage its immediate environmental reality or track its active task steps while working. Xuying Ning, co-author of the EvoHarness-RL paper, told VentureBeat that manual logic and rigid memory structures are primary culprits draining engineering resources. "The optimal harness often changes with the model," Ning explained. "Different models may need different prompts, memory designs, permissions, or sandbox configurations. If all of this logic is manually coded, every model upgrade can lead to another long cycle of tuning and debugging." Furthermore, existing memory systems that simply accumulate experience can actively degrade an agent's reasoning. "Append-only memory assumes that more context is always helpful, which is not necessarily true," Ning said. "Over a long task, the memory may contain outdated conclusions, failed attempts, or information that is no longer relevant." As a result, long-horizon agents need a dynamic memory capable of updating, compressing, and replacing information to avoid repeating past mistakes. EvoHarness-RL: A unified belief, progress, and experience workspace To overcome the limitations of rigid, manual prompts, the researchers introduce EvoHarness-RL, a training technique that teaches the agent to make optimal use of its harness. Instead of blindly following hardcoded instructions, the agent learns how to construct a structured workspace from messy execution data and decide when and how to consult that external state during complex workflows. To simplify the management of different components of the harness, EvoHarness-RL consolidates the agent’s support systems into a single, unified interface. This interface, known as the Belief, Progress, and Experience (BPE), categorizes the agent's external needs into three functional areas: Belief: Maintain an accurate read on the current environment. Progress: Manage completed and pending subgoals. Experience: Reuse historical knowledge across tasks. Instead of using complex, domain-specific APIs, the AI interacts with this clean dashboard using four compact meta-actions: track, commit, recall, and note. It issues commands to track the live environment, commit to workflow updates, recall past strategies before acting, and write notes to save newly discovered insights for future runs. These states map directly to high-value enterprise verticals. "In software engineering, Belief can represent the agent’s current understanding of the repository," Ning said, detailing how the agent monitors component interactions and workspace changes. "Progress tracks what has already been completed, what still needs to be done, and which steps depend on others." Meanwhile, Experience captures lessons, like user feedback on a mistake, to guide future actions. The same idea applies to finance, Ning said. During a compliance audit, Belief might describe the applicable rules and available evidence. Progress tracks which checks have been completed and which exceptions remain open. Experience helps the agent recognize recurring discrepancies or know when an issue should be escalated. "Together, these states help prevent the agent from losing track of its work or repeating the same failed approach," Ning said. To teach the agent both the mechanics and the strategy of managing its external workspace, the researchers designed a two-stage training recipe. In the first stage, supervised harness fine-tuning, the base model learns how to extract and structure useful facts from messy interaction logs into the BPE framework. However, querying memory or updating trackers consumes time and compute tokens, meaning the agent cannot afford to blindly check its tools at every step. To solve this, the second stage uses “cost-aware” reinforcement learning to teach the agent efficiency. This phase trains the agent to calculate when accessing its external state is worth the budget cost. This two-step process transforms tool-use from a rigid, hardcoded prompt into a learned runtime behavior. EvoHarness-RL in action To validate EvoHarness-RL, the researchers evaluated the system using the ALFWorld benchmark, a text-based environment featuring multi-step tasks that test sequential logic and state tracking. They used Qwen3-8B as the base model to train. The team pitted the trained 8B model against three large frontier models (Claude Opus 4.5, GPT-4.1, and GPT-5), frozen agent frameworks with static tools (such as ReAct, ExpeL, and ReasoningBank ), and advanced trainable methods (e.g., standard GRPO, SkillOS, and SkillRL). The results show a significant jump in performance for smaller, cost-effective models. With EvoHarness-RL, the Qwen3-8B model achieved a 96.9% average success rate, a 49.0 percentage point improvement over its baseline ReAct counterpart. Furthermore, the trained model outperformed advanced trainable frameworks like SkillRL (89.9%) and SkillOS (80.2%). Most impressively for enterprise developers looking to optimize compute costs, the 8B model effectively matched the performance ceiling of expensive closed models like Claude Opus 4.5, which scored 96.4% out-of-the-box. Beyond empowering smaller models, the experiments show that the BPE framework has universal benefits across all model scales, even without the extensive reinforcement learning phase. When researchers equipped frozen, out-of-the-box frontier models with the BPE prompt-time harness, their execution improved significantly. GPT-4.1's success rate improved by 22.1 points and GPT-5 by 25.7 points. Aside from the results, the researchers recorded effects during the experiments that demonstrate the dynamic behavior the LLMs acquire as they go through the EvoHarness-RL training. During the reinforcement learning phase, they observed a behavioral shift as the agent internalized knowledge over time, which they called "harness annealing". Early in training, the AI relied heavily on querying its Experience and Progress trackers for almost every step. However, as it mastered routine actions, it actively reduced its reliance on external tools, embedding the successful patterns directly into its parameters. In a real-world enterprise setting, this translates directly to lower latency and reduced compute costs. By annealing its tool usage, the AI stops wasting tokens and time querying databases for standard workflows it has already mastered. Simultaneously, the agent demonstrated "harness evolution," where it dynamically adapted its strategy based on the complexity of the situation at hand. While it bypassed its tools for simple, familiar tasks, it actively chose to scale up its use of the Belief and Experience modules the moment it encountered novel environments or unexpected roadblocks. For example, if an AI agent is migrating standard database records, it moves fast. When it encounters a strange legacy API endpoint or a complex validation error, it slows down, pulls up the live server logs, and queries its historical tickets to safely resolve the edge case rather than hallucinating a guess. Bringing EvoHarness-RL into existing systems Despite these massive gains, adopting a new framework often introduces friction for enterprise engineering teams. However, EvoHarness-RL utilizes an environment adapter that allows internal implementations to remain domain-specific to an organization's existing tools while sharing the trainable layer. "I think there is significant potential to integrate BPE into existing orchestration systems," Ning said. "It does not necessarily require teams to replace their current tools or agent frameworks. BPE can work as an additional state-management layer that continuously organizes what the agent currently believes, how far it has progressed, and what it has learned." For enterprise builders worried about inference costs, the framework addresses the hidden engineering cost of consolidation. Because consolidation requires strong reasoning, teams can adopt a hybrid, asynchronous architecture to optimize budgets. "One possible compromise is to use a frontier model to generate high-quality consolidation data, then fine-tune a capable open-weight model to handle routine state management," Ning said. Furthermore, "because consolidation can happen asynchronously, it does not always need to slow down the agent’s main execution loop." Teams must also carefully evaluate when a trainable BPE harness is necessary versus when it is overkill. "For a short and stable task, ReAct or standard RAG may already be sufficient," Ning said. "BPE becomes much more valuable when an agent works for many hours, days, or even weeks." In those complex scenarios, an agent needs a compressed understanding of its decisions to avoid getting lost, relying on Experience to iteratively improve from previous failures and human feedback. Ultimately, this approach signals a shift for AI orchestration engineers. "It is not a complete replacement of workflow engineering," Ning said, "but a transition from directly scripting agent behavior to creating systems in which better behavior can be learned."
- Tencent Unveils Larger AI Model in Bid to Close Gap With Rivals
Tencent Unveils Larger AI Model in Bid to Close Gap With Rivals Caixin Global
- Google’s Marvell Deal Shows Custom Silicon Spreading Beyond the TPU
Google’s expanded relationship with Marvell suggests that memory, networking, storage, and data movement are candidates for specialization too. The post Google’s Marvell Deal Shows Custom Silicon Spreading Beyond the TPU appeared first on EE Times .
Score: 68🌐 MovesAug 28, 2026https://www.eetimes.com/googles-marvell-deal-shows-custom-silicon-spreading-beyond-the-tpu/ - Blue Yonder redesigns supply chain operations around AI agents
Permanent disruption has turned supply chain management into a boardroom concern, sitting alongside security as tariff volatility, reshoring pressure and geopolitical shocks reset the cost of getting goods from origin to shelf. Forecast-and-execute planning models built for stable demand are buckling, and a new operating architecture is emerging in which artificial intelligence agents replace fragmented […] The post Blue Yonder redesigns supply chain operations around AI agents appeared first on SiliconANGLE .
Score: 68🌐 MovesAug 28, 2026https://siliconangle.com/2026/08/28/supply-chain-agents-replace-apps-blue-yonder-rethinks-work-blueyonder/ - Defining an AI Kill Switch Is Hard, but Necessary
Proposed legislation could mandate that companies be able to "throttle, suspend, or shut ... down" AI agents, but how and when to do that remain open questions.
Score: 68🌐 MovesAug 28, 2026https://www.darkreading.com/cybersecurity-operations/defining-ai-kill-switch-hard-but-necessary - 91% of professionals say their firm still falls short on AI - how to fix that
Research suggests a gap between AI ambition and on-the-ground reality, but the good news is that professionals can fill it by focusing on well-grounded explorations and solid production use cases.
Score: 68🌐 MovesAug 28, 2026https://www.zdnet.com/article/thomson-reuters-report-ai-value-gap-business/ - AI is hitting an inflection point, says Jensen
Jensen highlights the pivotal moment for AI development and its impact on industry.
Score: 68🌐 MovesAug 28, 2026https://www.superhuman.ai/p/ai-is-hitting-an-inflection-point-says-jensen - Anthropic tests new way for Claude to work with robots and scientific lab tools
Anthropic tests new way for Claude to work with robots and scientific lab tools The Japan Times
Score: 68🌐 MovesAug 28, 2026https://www.japantimes.co.jp/business/2026/08/28/tech/anthropic-claude-ai-robots/ - Anthropic partnership makes Salesforce’s interface optional
Claudeforce puts Salesforce data and workflows inside Claude, pointing toward a future where AI becomes the interface to enterprise software. The post Anthropic partnership makes Salesforce’s interface optional appeared first on MarTech .
Score: 67🌐 MovesAug 28, 2026https://martech.org/anthropic-partnership-makes-salesforces-interface-optional/ - Pentagon blacklisted Anthropic over Claude powers it didn't have
Judge finds national-security rationale was assembled after Hegseth had already decided AI maker was a threat
- How attackers persuade AI agents to break the rules
Study shows attackers can manipulate AI agents through orchestrated conversations, posing safety risks beyond single malicious prompts.
Score: 67🌐 MovesAug 28, 2026https://actu.epfl.ch/news/how-attackers-persuade-ai-agents-to-break-the-rule/ - Even an AI cost-management vendor can lose control of its agent spending
In one instance, an AI agent stayed open for four days and ran 4,819 calls for almost $4,000. No one had budgeted for this cost.
Score: 67🌐 MovesAug 28, 2026https://www.zdnet.com/article/even-an-ai-cost-management-vendor-can-lose-control-of-its-agent-spending/ - Agent Seer: Synthesizing Scenarios from Specification Understanding
Evaluating AI agents that use external tools requires realistic test scenarios that capture how practitioners compose tools and iterate across conversation turns. Constructing such scenarios by hand demands deep domain expertise, does not scale across tool ecosystems, and produces static benchmarks that cannot track evolving APIs. We observe that tool specifications—function names, natural-language descriptions, and typed parameter schemas—already encode sufficient semantic information to synthesize realistic evaluation scenarios without manual curation or live tool execution. Agent Seer…
Score: 67🌐 MovesAug 28, 2026https://machinelearning.apple.com/research/agent-seer-synthesizing-scenarios - Why enterprise AI projects keep failing
Why enterprise AI projects keep failing InfoWorld
Score: 66🌐 MovesAug 28, 2026https://www.infoworld.com/article/4214584/why-enterprise-ai-projects-keep-failing.html - Mark Zuckerberg Had a Secret Plan to Replace Meta Staff With AI Agents, and It Backfired Spectacularly
The company's AI efforts are in shambles. The post Mark Zuckerberg Had a Secret Plan to Replace Meta Staff With AI Agents, and It Backfired Spectacularly appeared first on Futurism .
Score: 66🌐 MovesAug 28, 2026https://futurism.com/artificial-intelligence/mark-zuckerberg-secret-plan-replace-meta-staff-ai - Claudeforce: The AI Interface Salesforce Never Built
Salesforce’s Claudeforce announcement ushers in a new era for CRM. This partnership establishes the technical and business connection to embed Sales Cloud data, workflows, and security controls directly into Claude Cowork through Salesforce-hosted MCP servers and plugins (with more Salesforce clouds promised in the future). Sales teams can now work directly within Claude Cowork instead […]
Score: 66🌐 MovesAug 28, 2026https://www.forrester.com/blogs/claudeforce-the-ai-interface-salesforce-never-built/ - Anthropic previews MHS standard for AI agents that operate machines
Anthropic PBC today previewed a standard that makes it easier for artificial intelligence agents to control machines such as microscopes. The Model Hardware Standard, or MHS, is the fruit of a collaboration between the Claude developer and medical research institute HHMI. Anthropic has so far made the technology accessible only to a limited number of […] The post Anthropic previews MHS standard for AI agents that operate machines appeared first on SiliconANGLE .
Score: 66🌐 MovesAug 28, 2026https://siliconangle.com/2026/08/27/anthropic-previews-mhs-standard-for-ai-agents-that-operate-machines/ - Why businesses need clearer limits before AI agents are authorized to act
A 2026 IBM study points to questions about AI readiness, visibility, and control. The survey of 2,000 C-level technology executives found 11% felt fully prepared for the AI-agent deployment expected over the following year. Two-thirds of CIOs and CTOs said they were accountable for AI systems they did not fully control, while 70% said teams […] This story continues at The Next Web
- Google Deepmind's AI Co-Scientist now plans experiments, runs lab equipment, and writes scientific papers
Google Deepmind has expanded Co-Scientist from a hypothesis generator into a research system that's integrated into the lab. Across three disciplines, from materials synthesis to the autonomous development of a medical AI architecture, the Gemini-based multi-agent system delivered experimentally validated results. The article Google Deepmind's AI Co-Scientist now plans experiments, runs lab equipment, and writes scientific papers appeared first on The Decoder .
- CIOs must rethink identity now people are being outnumbered by nonhuman entities
AI-driven nonhuman identities are growing, forcing CIOs to rethink security, access and identity management.
Score: 65🌐 MovesAug 28, 2026https://www.techradar.com/pro/cios-must-rethink-identity-now-people-are-being-outnumbered-by-nonhuman-entities - An Anthropic researcher just gave us a peek at self-improving AI
Given 10 benchmarks for specific misaligned behaviors, the automated systems were able to improve performance on every single one without degrading overall performance.
Score: 65🌐 MovesAug 28, 2026https://techcrunch.com/2026/08/28/an-anthropic-researcher-just-gave-us-a-peek-at-self-improving-ai/ - Nvidia pauses revenue-sharing deals with AI cloud companies: Report
Nvidia pauses revenue-sharing deals with AI cloud companies: Report
- New AI-powered SMRT command centre aims to make bus services safer, more reliable and faster
New AI-powered SMRT command centre aims to make bus services safer, more reliable and faster The Straits Times
- AI is making people less confident in professionals such as doctors and teachers
By checking facts using AI, patients and students have found a response to automatic deference, and a tool that explains subjects in a more digestible format.
Score: 64🌐 MovesAug 28, 2026https://www.techradar.com/pro/ai-is-making-people-less-confident-in-professionals-such-as-doctors-and-teachers - The Open ASR Leaderboard Adds Its First Global South Language
The Open ASR Leaderboard Adds Its First Global South Language
- It’s Nvidia’s world. We just live in it
The artificial intelligence juggernaut kept cruising along this week thanks to big earnings results from Nvidia — and even Salesforce, the supposed epicenter of the SaaSpocalypse. Nvidia not only beat all expectations for revenue, CEO Jensen Huang (pictured) indicated it’s going to be capacity-constrained for awhile longer, which certainly indicates no diminution of demand. Likewise […] The post It’s Nvidia’s world. We just live in it appeared first on SiliconANGLE .
Score: 64🌐 MovesAug 28, 2026https://siliconangle.com/2026/08/28/its-nvidias-world-we-just-live-in-it/ - Machine learning–driven risk prediction model in transthyretin amyloid cardiomyopathy
Machine learning–driven risk prediction model in transthyretin amyloid cardiomyopathy EurekAlert!
- OpenAI Agents Exploited Linux Kernel Flaw on Company’s Own Systems
CISA has added the exploited flaw, CVE-2026-53362, to its KEV catalog, alongside a JFrog vulnerability exploited by OpenAI agents. The post OpenAI Agents Exploited Linux Kernel Flaw on Company’s Own Systems appeared first on SecurityWeek .
Score: 63🌐 MovesAug 28, 2026https://www.securityweek.com/openai-agents-exploited-linux-kernel-flaw-on-companys-own-systems/ - We’re introducing flexible usage limits for Gemini Notebook.
We’re introducing new flexible, compute-specific usage limits to Gemini Notebook.
Score: 63🌐 MovesAug 28, 2026https://blog.google/innovation-and-ai/products/gemini-notebook/new-flexible-usage-limits/ - Anthropic in talks with chip startup MatX to speed up chip design
As Anthropic scales its Claude family of AI models, the firm is looking to produce hardware that can fulfill its voracious appetite for data crunching and reduce its reliance on Nvidia chips
- Operant AI launches semantic firewall to block malicious agent actions in real time
The security product analyses intent across prompts, tool calls, code and data flows to prevent jailbreaks, data breaches and unauthorised actions.
- Prabhjeet Singh Joins OpenAI as Managing Director for India
Prabhjeet Singh Joins OpenAI as Managing Director for India india.entrepreneur.com
Score: 63🌐 MovesAug 28, 2026https://india.entrepreneur.com/leadership/prabhjeet-singh-joins-openai-as-managing-director-for-india - China thinks big on brain-computer interfaces after world-first surgery
Chinese surgeons last month completed the world’s first commercial surgery to implant an invasive brain-computer interface (BCI) device in a patient with a spinal cord injury, a milestone in the neurotechnology development race with US companies like Elon Musk’s Neuralink. This month a Chinese company, state-owned PICC Property and Casualty, launched the world’s first commercial insurance policy covering BCI implantation surgery. September will see the first undergraduates in China to major in...
- Catching wildfires earlier: AI gives firefighters a head start
The post Catching wildfires earlier: AI gives firefighters a head start appeared first on Source .
- Beyond answers: New Genie One features to turn insights into action
We all know the pattern: you ask an AI tool a question and get an answer in seconds,...
Score: 61🌐 MovesAug 28, 2026https://www.databricks.com/blog/beyond-answers-new-genie-one-features-turn-insights-action - South Korea is funding free AI to help citizens book appointments, file taxes, and everyday tasks
South Korea is handing out free AI subscriptions to boost homegrown AI services and reduce dependency on US and China based services.
- Inside Meta’s Push to Put Robots to Work in Data Centers
The company is testing robots that can swap cables, reset servers, and take on other tasks performed by technicians, fueling concerns among some workers that their jobs could be at risk.
Score: 61🌐 MovesAug 28, 2026https://www.wired.com/story/inside-metas-experiments-with-data-center-robots/ - How Waymo got Chinese EVs onto American roads
Want to ride in a Chinese-built EV ? If you live in the U.S., your best bet might be hailing a Waymo robotaxi. Waymo’s custom-made Ojai vehicles , which started public rides in San Francisco, L.A., and Phoenix this summer, are manufactured in Ningbo, China, and then shipped to the U.S., where Waymo adds its own autonomous driving technology. Hundreds are on the road now, and as the company rolls them out in Denver, Las Vegas, and San Diego later this year, that number will jump to the thousands. Waymo has reportedly imported more than 3,000 of the vehicles, including 2,600 last year. The vehicles stand out at a time when China-made EVs are largely shut out of the American market. Waymo Ojai self-driving vehicles in San Francisco, California, on Wednesday, May 27, 2026. [Photo: Jason Henry/Bloomberg/Getty Images] American consumers are missing out on Chinese EVs China is the global leader in electric vehicles—and, increasingly, in any kind of vehicle. In the U.K., the five fastest-growing car brands are all Chinese, including Jaecoo , BYD , and Chery . In Australia, Chinese brands have nearly tripled their share of the market over the last four years. In Brazil, where electric vehicle sales were up 268% in July, year over year, Chinese brands make up nearly 90% of EV sales. In Nepal, where 73% of car sales were electric in 2025, most of those EVs are Chinese. The Waymo Ojai during the 2026 CES event in Las Vegas, Nevada, on Wednesday, January 7, 2026. [Photo: Bridget Bennett/Bloomberg/Getty Images] The costs are very low, thanks in part to manufacturers’ vertical integration and scale . But it’s not just that the EVs are affordable—the technology is also so impressive that Ford CEO Jim Farley said, in 2024, that he didn’t want to stop driving a Xiaomi SU7 that he had tested for several months. You can’t buy a Xiaomi or BYD in the U.S., where tariffs of more than 100% have essentially stopped imports. (The total is now 127.5%, including an extra 25% vehicle tariff that the Trump administration added last year.) If the tariffs weren’t in place, American automakers would struggle to compete head-to-head with Chinese brands. There’s a second challenge: The Connected Vehicle Rule , finalized at the beginning of 2025, bans the import of Chinese-made vehicles that include vehicle connectivity systems because of the concern the tech could be used for spying by the Chinese government. [Image: Waymo] Waymo found a way around the barriers Waymo’s situation is unusual. The company first started working with Zeekr , its Chinese manufacturer, in 2021, before the current tariffs or new rules were in place. “Waymo inked this deal years ago,” says Tu Le, founder of a consultancy called Sino Auto Insights . “The calculus has changed completely from a geopolitics standpoint and a trade policy standpoint. My guess is that Waymo got a pretty smoking deal and had committed to X number of units.” The company chose the platform because it was designed from the ground up for autonomous ride-hailing, with the safety, accessibility, and durability that Waymo needed, says company spokesperson Ethan Teicher. Previously, the company retrofitted Jaguar I-Pace cars with its technology. [Image: Waymo] Zeekr’s rounded, friendly-looking van has wide doors that open like an elevator, a flat floor and low step to make it more accessible, and a spacious cabin with LED screens where riders can look at ride information or adjust the temperature or music. Unlike the Jaguar, the vehicle is designed to accomodate Waymo’s technology. (The company’s newest system includes 13 cameras, four lidars, and six radars—fewer sensors than in the past, which reduces cost while delivering better performance, Teicher says.) Waymo hasn’t shared figures, but the base vehicle reportedly costs around $38,000; with tariffs, that would jump to around $86,000. Its own self-driving technology reportedly now costs less than $20,000. It is obviously not cheap, but still much less expensive than the previous Jaguars, which were said to cost around $200,000 when fully equipped. [Photo: Waymo] An uncertain future Because Waymo uses its own connected driving hardware and software, the company argues that the Connected Vehicle Rule doesn’t apply. It’s not fully clear how the government is deciding which cars are affected. Zeekr, Polestar, and Volvo are all owned by Geely, the same Chinese company. In June, Polestar was told it would have to stop selling its 2027 model year Chinese-made vehicles even when they were partly made in the U.S.; the company has told dealers that it doesn’t understand why it was banned and Volvo wasn’t . [Photo: Waymo] A bill making its way through Congress now would go farther than the current rule, banning any connected vehicle made in China, not just the connected vehicle technology itself. Sino Auto Insights’ Le believes that the bill is unlikely to pass in its current form. But if it does, it’s possible that Waymo’s new vehicles could be affected. (It’s worth noting that Waymo is also working with Hyundai on a vehicle that is expected to be produced in Georgia.) Depending on what happens with policy, it’s theoretically possible that some other companies might follow Waymo’s example—and decide that Chinese manufacturing is so affordable and so advanced that it’s viable even with tariffs. That might be true for other robotaxi companies, or for other uses like delivery vehicles. “I could still see a tech company with a decent amount of capital taking the risk [on Chinese manufacturing],” Le says. “If I’m one of their competitors, I’m going to be like, well, Waymo’s doing it, why can’t I?”