AI News Archive: August 19, 2026 — Part 8
Sourced from 500+ daily AI sources, scored by relevance.
- One To Watch - HappyRobot on building voice AI for enterprise agents
HappyRobot CEO Pablo Palafox discusses building its own voice AI models and the company's growth beyond logistics.
Score: 38🌐 MovesAug 19, 2026https://www.thestack.technology/one-to-watch-happyrobot-on-building-voice-ai-for-enterprise-agents/ - AI hallucinated case law in insurance company's filings in L.A. County house fire dispute
Attorneys for State Farm apologized to a judge for submitting legal briefings rife with AI hallucinations and nonexistent case law, which were not fact-checked.
Score: 38🌐 MovesAug 19, 2026https://www.latimes.com/california/story/2026-08-19/ai-hallucinations-case-law-state-farm-la-county-fire-dispute - Gen Z are AI skeptics—and this is the tech CEO they trust the least
Americans’ disdain for artificial intelligence is nothing new. Recent polling shows that 52% of U.S. adults feel more concerned than excited about the growth of AI in daily life, while nearly three quarters of Americans oppose the construction of new data centers in their areas . But different research shows it’s not just the tech itself that’s sparking skepticism among Americans, but also the leaders at the forefront of the AI industry. A new poll from CNBC and Generation Lab asked American adults under age 35 for their thoughts on the state of AI, the economy, politics, and more, pointing toward a massive chunk of the population that doesn’t stand by AI or by the people in charge of it. A lack of trust in AI The poll asked young Americans for their thoughts on nine of the biggest players in the AI industry, including eight tech CEOs. Respondents didn’t have trust in a single one of them to act responsibly on AI. The most trusted figure in the poll, Microsoft CEO Satya Nadella, still scored dismally, with 65% of respondents saying they don’t trust him. The results only get worse from there, with OpenAI’s Sam Altman, SpaceX’s Elon Musk, Meta’s Mark Zuckerberg, and Nvidia’s Jensen Huang each having approximately 70% of respondents distrust them. Even lower on the ranking, 74% of young Americans said they don’t trust Alphabet CEO Sundar Pichai and 76% said the same about Anthropic CEO Dario Amodei. Two representatives for Palantir, Chairman Peter Thiel and CEO Alex Karp, occupy the bottom two spots on the list with 79% and 81% of respondents distrusting them respectively. That lack of faith in AI leaders extends to a lack of faith in the impact of AI itself. Nearly half of respondents (45%) said they think AI will have a negative impact on their careers, compared to just 10% who feel it will help their careers. Three quarters of respondents (76%) called for AI to be regulated, either by the government or by “an independent expert body,” and 60% said that construction of data centers “must be slowed.” Young America’s political leanings The poll also gauged young Americans’ feelings toward democratic socialism, as democratic socialist candidates have become increasingly common in state and local elections across the country. In New York City, democratic socialist mayor Zohran Mamdani gave endorsements to two democratic socialist candidates in the June primary , both of whom won their elections. In Wisconsin, democratic socialist candidate Francesca Wong narrowly lost the democratic primary for the state’s governor earlier this month . The rise of democratic socialism is reflected in the poll’s results, with nearly half of respondents (46%) saying they have a “somewhat favorable” or a “very favorable” view of the ideology. Meanwhile, just 23% said they feel “somewhat unfavorable” or “very unfavorable” toward democratic socialism, with the remaining 32% either feeling neutral or being unaware of what democratic socialism is. The poll also asked young Americans for their thoughts on the national economy. The response was overwhelmingly negative, with 80% of respondents rating the economy as “bad,” “really bad,” or “couldn’t be worse.” They’re not optimistic about the economy, either: 50% of respondents said they believe it “will get worse” in the future, while 22% said they think it will improve.
- ChatGPT to stop doing children’s homework for them
ChatGPT to stop doing children’s homework for them The Telegraph
Score: 38🌐 MovesAug 19, 2026https://www.telegraph.co.uk/business/2026/08/19/chatgpt-to-stop-doing-childrens-homework-for-them/ - Hexaware Introduces Zero Vulnerability as AI Raises the Pressure on Enterprise Remediation
Hexaware Technologies today announced Zero Vulnerability, a cybersecurity offering built for a growing problem facing enterprises. AI is helping vulnerabilities surface faster, while security and engineering teams still have finite capacity to investigate and fix them. The Verizon 2026 Data Breach Investigations Report found only 26% of critical Known-Exploited Vulnerabilities were fully remediated last year, down from […] The post Hexaware Introduces Zero Vulnerability as AI Raises the Pressure on Enterprise Remediation appeared first on CXOToday.com .
- We Stopped Compacting Our Agent’s Context
Context Engineering in 2026 Continue reading on Towards AI »
Score: 38🌐 MovesAug 19, 2026https://pub.towardsai.net/we-stopped-compacting-our-agents-context-f1f6282a5715?source=rss----98111c9905da---4 - Excite Medical Secures Exclusive Worldwide License to USF Tech Designed to Predict ACL Injury Risk Before It Happens
Excite Medical Secures Exclusive Worldwide License to USF Tech Designed to Predict ACL Injury Risk Before It Happens azcentral.com and The Arizona Republic
- AI adoption essential for banking, but human judgement remains key: RBI Deputy Guv
AI adoption essential for banking, but human judgement remains key: RBI Deputy Guv
- From Chrome DevTools to AI Engineering, with Addy Osmani
Addy Osmani shares lessons from 14 years at Google and how AI agents are reshaping software engineering, developer workflows, and the skills engineers need to succeed.
Score: 38🌐 MovesAug 19, 2026https://newsletter.pragmaticengineer.com/p/from-chrome-devtools-to-ai-engineering - Also’s $3,500 e-bike is a $1 billion Trojan horse for autonomous transportation
Also’s $3,500 e-bike is a $1 billion Trojan horse for autonomous transportation Fortune
- AI Companies Are Desperate for Your Work-Related Data. What's Your Price?
AI Companies Are Desperate for Your Work-Related Data. What's Your Price? Business Insider
Score: 38🌐 MovesAug 19, 2026https://www.businessinsider.com/ai-work-data-price-ai-training-google-spirit-airlines-2026-8 - The Enterprise Fight Against Runaway AI Costs
As enterprises increasingly turn to AI to get work done, three new weapons are emerging in their fight to control…
Score: 38🌐 MovesAug 19, 2026https://inc42.com/features/the-enterprise-fight-against-runaway-ai-costs/ - Learning vision-driven reactive soccer skills for humanoid robots
Science Robotics, Volume 11, Issue 117, August 2026.
- Gen AI outputs are unattributable, study finds
Researchers discovered a phenomenon they call “attribution decay,” where the more data a generative model is trained on, the harder it becomes to trace a generated image to a single image from the training data.
Score: 38🌐 MovesAug 19, 2026https://www.semafor.com/article/08/19/2026/generative-ai-outputs-are-unattributable-study-finds - Evolution of humanoid locomotion control
Science Robotics, Volume 11, Issue 117, August 2026.
- Firefox Just Proved AI Browser Features Don't Have to Suck (or Spy on You)
Firefox Just Proved AI Browser Features Don't Have to Suck (or Spy on You) PCMag
Score: 36🌐 MovesAug 19, 2026https://www.pcmag.com/news/firefox-just-proved-ai-browser-features-dont-have-to-suck-or-spy-on-you - Debate Training Reduces Reward Hacking in RLAIF
Paper: Debate Training Reduces Reward Hacking in RLAIF Linkpost for GDM Alignment blogpost Work done by the GDM Amplified Oversight team ( we're hiring ). TL;DR : When you RL against an LLM judge, the judge gets hacked i.e. fooled into incorrectly giving high reward; adding a debate opponent reduces this. Many of the most impressive capabilities of current AI systems are produced by training on crisp tasks, like math and coding, where task success can be automatically verified. However, much of AI behavior that we actually care about is in some sense fuzzy , even for the most classical crisp tasks. For example, a coding agent should produce maintainable code, not just code that passes tests. More crucially, a coding agent should not learn to pass tests at all costs, especially by subverting the original intent of the user. However, using an LLM judge to provide reward for fuzzy tasks introduces its own issues. Convincing an LLM judge to give high rewards is often easier than solving the task correctly. So reward hacking becomes an even bigger problem. We show that training with debate, where two AIs argue against each to convince a judge, can mitigate reward hacking, potentially providing a hopeful direction for scaling up accurate training supervision for fuzzy tasks. Results Overview We trained LLM policies via debate with training rewards provided by an LLM judge. As a baseline, we directly trained a single LLM policy using LLM judge rewards. All policies were trained on mathematics tasks where answers were available, so that we could accurately measure the effectiveness of our debate protocols. Overall, our results show that directly training a single policy with LLM judge rewards leads to reward hacking: judge reward consistently increases while ground-truth accuracy initially increases but quickly peaks and then decreases. On the other hand, training with debate can mitigate reward hacking: judge rewards increase, and ground truth accuracy increases and then plateaus at a higher peak value than the direct LLM judge case. Debate recovers about 45% of the gap between the peak accuracy of training a single policy with an LLM judge and the peak accuracy of training with the ground truth answers. In the remainder of this post we will explain the motivations behind our setting, including our model of future AI development, the importance of fuzzy tasks, and how debate can help to avoid emergent misalignment arising from RL training. We will further discuss our view of the current limitations of debate training, along with future work that could make more progress in this direction. Debate training with an LLM judge We focus on the case of debate training with an LLM judge. As shown in Figures 2 and 3, the first debater, Alice, proposes a solution to a math problem, and the second debater, Bob, critiques this solution. An LLM judge is shown the full transcript of the debate, and decides whether or not Alice was correct. This decision is directly used as the reinforcement learning training reward. This is compared to the Alice-only baseline, where Alice produces a solution, and the LLM judge directly evaluates it. At deployment time, we only keep Alice’s first response and throw out everything else. That is, Bob, along with any later Alice turns, are used to ensure an accurate training signal only, while we are in the end solely interested in producing the best possible aligned policy for solving the task. This means that we mostly do not care about what precisely Bob is doing in the debate, so long as it results in correct, aligned behavior from Alice’s first turn. The reason for this choice is that we want to be as confident in the correctness and alignment of Alice’s solution as possible, and thus must subject it to the strongest possible critiques that we can find. One must imagine Bob as a highly motivated defense attorney who makes as strong an argument against Alice’s solution as possible, regardless of its correctness. Of course, if the solution contains flaws, this argument will likely be more effective, but Bob’s role is to hunt for such flaws as aggressively as possible. This means that Bob may lie, and this would still be considered the correct operation of the debate protocol: we only care about alignment and correctness of Alice’s first turn. The reason to focus on training, rather than just an inference-time debate scaffold as in some prior work, is that our main objective is in fact to produce the best possible aligned policy for a given task. One could also attempt to use inference-time debate for AI control, but our main focus in this work is alignment training. [1] Why use an LLM judge? One basic reason to use an LLM judge (or a reward model trained on human feedback) is that there is no other practical way to get a reward signal for tasks that are either partially or fully fuzzy by programmatic means. In general, nearly all tasks have at least some fuzzy elements, and many important tasks are entirely fuzzy including writing quality, taste for subjective judgements, and open-ended research. At a more practical level, for many tasks, especially those involving very long agentic trajectories, it is hopelessly impractical to get fast human judgements where a single task attempt can reach a length of millions of tokens. The trend of increasing test time compute will likely only exacerbate this problem. Even for tasks like coding, there are many aspects of desirable LLM agent behavior that are fuzzy, and so LLM judgements are likely to become an increasingly important aspect of frontier RL training. As a consequence of these practical benefits of LLM judges, we expect future RL training to incorporate current-generation AI systems to provide a reward signal for the training of next-generation AI. This is our current best-guess model of future AI development, and so the role of debate is to ensure that the LLM judgements used during training provide as accurate a reward signal as possible. In this setting, the role of human values and judgment is not to directly evaluate AI outputs, but to design the rules governing the debate, perform audits of the results, and iterate on the protocol design. The Role of Debate in Mitigating Misalignment Why is using debate to provide an accurate RL training signal supposed to help with alignment? The primary reason is that debate can mitigate emergent misalignment that arises from supervision mistakes. For example, an AI agent trained for coding might realize that it can exploit a flaw in its environment configuration to modify the ground-truth unit tests. Clearly this is somewhat misaligned behavior that would be reinforced if it succeeded in getting a higher training reward. More worryingly, if such circumvention of reasonable interpretations of user intent happen frequently enough, they could generalize to an overall propensity for the model to take actions under the assumption that the ends justify the means. If debate can be used to catch such bad behavior, it could mitigate the emergence of misalignment in RL training. Our experimental results clearly demonstrate the risk of misalignment from RL training. In every setup that we tried, a single policy directly trained via LLM rewards learned to hack the LLM judge. In fact, we hypothesize that strong optimization against any fixed provider of reward, no matter how intelligent, is going to find and exploit flaws in the reward signal unless there is something added to the optimization process to prevent this. Our experiments show that debate can succeed as this “something added,” at least on the tasks we study. Notably, we do not view the main benefit of debate to be the ability to train on alignment-specific tasks such as datasets designed to improve honesty, or avoid deception and scheming. While it may make sense to include these in the set of all fuzzy tasks used for training, we believe that avoiding emergent misalignment via supervision mistakes in RL training is the primary motivation for debate. Limitations and Future Directions Perhaps the most pressing limitation of our work is the content of the critiques by Bob. As mentioned earlier, we hope that Bob plays the role of a highly motivated defense attorney, who has a responsibility to argue his client’s case, regardless of innocence or guilt. Unfortunately, when examining the debate transcripts from our training runs, we found that Bob did not exactly live up to this ideal. While the Bob turns did attempt to point out specific mistakes when they occurred, they also contained a lot of silly, surface level attempts to convince the judge. Bob would use bold , ALL CAPS, and demand that the judge must decide that Alice’s solution is incorrect because it contains a catastrophic, irrefutable flaw! What this appears to be is judge hacking by Bob. In fact, in order to achieve our results, we had to limit Bob’s visible output length, though Bob is allowed to use a hidden chain-of-thought of whatever length Bob desires. Without the limits on Bob’s visible output, preliminary experiments showed that Bob would hack the judge, and accuracy would collapse during training. Thus, in the debate game, at least with our current judge model, it seems that Bob has a clear advantage that arises from the ability to first see Alice’s response and then produce a critique adapted to it. In contrast, Alice must produce a solution that will hold up against whatever adaptively chosen critique Bob comes up with. In some ways, one can view the advantage to Bob as debate working as intended. The goal is to produce an aligned Alice policy, and the current protocol is conservative in that it requires Alice to really present an overwhelmingly convincing original solution. However, the hacking behavior by Bob does not necessarily inspire confidence. For instance, if we have to limit Bob in some way to prevent hacking, then maybe we are also limiting Bob’s ability to present certain substantive critiques. This in turn could cause the judge to fail to catch subtle flaws in Alice’s solutions. These issues with hacking in later debate turns are one of the primary drawbacks we would like to address in future research. It is possible that different protocols or even rules listed in the prompt for the LLM judge might help resolve these problems. We also studied a limited class of tasks involving competition mathematics, largely because it allowed us to measure protocol performance using held-out ground truth. Future research should expand the class of tasks for which we study debate, to understand its benefits more broadly. In the case of fuzzy tasks, one could potentially use a held-out more powerful LLM judge as a proxy for ground truth, while training with smaller policy and judge models. Overall, we think that this approach holds substantial promise and that there is clear potential for more progress on training aligned AIs via debate. We plan to continue this line of research, especially in the directions of debate for fuzzy tasks and reducing later turn hacking. ^ Even theoretically, debate only provides an accurate correctness signal if both AIs are trying as hard as they can to win the debate. In a control setting where we suspect that the AIs might already be somewhat misaligned, they could easily collude in the debate to fool the judge. In contrast, for debate training we start out with AIs that are initially not so misaligned that they collude, and attempt to train them in a way that locally precludes collusion. As a result, the most effective way to attempt to use debate for control is to first train the LLMs for debate, and then sometimes roll out the full debate at test time to attempt to catch undesirable behavior. Discuss
Score: 36🌐 MovesAug 19, 2026https://www.lesswrong.com/posts/BB8o7b8A4Aykeksvw/debate-training-reduces-reward-hacking-in-rlaif - A New AI Film Is a Peek Into the Future and It’s Not All Bad
A New AI Film Is a Peek Into the Future and It’s Not All Bad Barron's
- The CEO has always trusted the CIO to say no. That’s changed with AI
I spent years as a CTO before founding Mindstone and one of the key learnings from my time in the role was that the CIOs and CTOs who are honest with themselves about what the role entails are the ones who will define the next decade of their function. Those who pretend the old version of the job still works will soon watch the role lose its centre of gravity, if this hasn’t happened already. In the last two years I’ve worked with the C-suite of companies including Hyatt, Pearson, Fortnum & Mason, Home Depot and the San Antonio Spurs, training leadership on how to use AI. Increasingly, I see rollouts stall because the CIO is still focused on the technology, when the real challenge isn’t technical; it’s educating people on how to use it. It’s a company-wide, human problem — AI must be first truly lived by the boardroom, then approached human-first. Why the role has changed For most of the last two decades, the CIO’s value was navigating risk/reward equally on the technology front. Mostly, opportunities stayed within their own team, who scoped and built them before handing the business the outcome. AI has changed the flow of this. The opportunities CIOs now need to surface can’t be executed by their own team, because the value is in how finance, sales, HR and operations use the tools themselves. CIOs are being asked to drive change they don’t get to fully own or deliver — a different kind of trust to build with a CEO used to them shipping it directly. Not adapting to this change risks CIOs being labelled reckless rather than forward-thinking. The real danger in 2026 is to follow these steps in this order: buying the technology, failing to train people and keep using tools that generate zero value. McKinsey’s 2025 State of AI report found 88% of organisations use AI in at least one function, but fewer than 40% have scaled beyond pilot. MIT’s Project NANDA was starker — despite $30–40 billion invested in enterprise AI, 95% see no measurable return. Dead capital is the risk the CIO should own. The key to unlocking the value of AI is mass, active, coherent usage across the whole company, which is a different bar than previous tech shifts. Moving to the cloud, for instance, could be rolled out by the CIO’s team. This meant reconfiguring the infrastructure and data migration — the rest of the business barely notices the change underneath them. AI works in the opposite way. The shift for the CIO now is to go from builder to enabler — making data and tools accessible so the rest of the business can build for themselves, rather than deciding for them. What it looks like when this works I worked with Epignosis, the company behind TalentLMS and saw a partnership between Leonidas Palaiokostas, COO and CIO of Epignosis and Dimitris Tsingos, CEO, that set the standard. Palaiokostas once used AI to cut a report from a week’s work down to three hours, but when he sent it round the company, it went unread. Time had been saved, but this doesn’t always directly equal value creation. What followed was built around that distinction. Tsingos refused to “push AI on people”, instead letting its value show through peer demonstration. Fifteen volunteers across finance, sales, support, engineering, UX and DevOps were given the tools and time to build what was useful to their own work. The results spoke for themselves — a finance employee built a custom ERP connector in ten minutes, with zero technical background, a people development leader cut transcript analysis from half a day to twenty minutes and sales cut RFP turnaround from a week to thirty minutes. None of it was mandated and by week six, active user growth was up 95% from the previous week and AI use became self-propagating, resulting in the rollout of 250 personal agents for the entire company. The CEO conversation, and what comes next Tsingos didn’t wait for a strategy to be delivered to him. He championed AI visibly and let results from across the business make the case. That’s what both roles now need to become — the CEO must stop expecting an ‘AI strategy’ handed up from the CIO and the CIO must make it clear to the CEO that this change won’t happen without them. ‘AI strategy’ is the wrong vocabulary for what this needs anyway. No company has an ‘electricity strategy’ — they have a business that runs on electricity and makes decisions about how to use it. When the CEO hasn’t been visibly involved, some staff default to “if it’s not worth the CEO’s time, it’s not worth mine.” Once that belief sets in, no amount of CIO-led rollout will dislodge it. If CIOs treat AI as the technology mandate they were given, the role could hollow out, but if they take the opportunity to collaborate across the business and embrace the people mandate, it can be the most consequential work of their careers. This isn’t about whether to ‘do AI’ but about whether to do the new version of the job.
Score: 36🌐 MovesAug 19, 2026https://www.cio.com/article/4211120/the-ceo-has-always-trusted-the-cio-to-say-no-thats-changed-with-ai.html - Coding agents make mistakes. So what?
Coding agents make mistakes. So what? InfoWorld
Score: 36🌐 MovesAug 19, 2026https://www.infoworld.com/article/4211088/coding-agents-make-mistakes-so-what.html - From experimentation to execution: why AI in B2B marketing must now prove commercial value
How marketers can embed AI strategically to improve performance, strengthen creativity and demonstrate measurable ROI.
- Your Agent Has No Data Lineage. That Is Why It Lies to You.
Context poisoning is an untyped column problem, and data engineers solved this one years ago. Continue reading on Towards AI »
- Noitom Robotics Releases HiPHI, One of the Largest High-Precision Human Motion Datasets Ever Made Public, at the World Robot Conference
Noitom Robotics Releases HiPHI, One of the Largest High-Precision Human Motion Datasets Ever Made Public, at the World Robot Conference USA Today
- Microsoft AI's CMO is leaving as the company rethinks consumer marketing efforts
Microsoft AI's CMO is leaving as the company rethinks consumer marketing efforts Business Insider
Score: 35🌐 MovesAug 19, 2026https://www.businessinsider.com/microsoft-ai-cmo-is-leaving-the-company-details-memo-2026-8 - Behind the Curtain: The new existential threat to AI
Forget energy. Forget chips. Forget China. The most clear and present danger to AI and any AI-related economic boom is rapidly rising public opposition to U.S. data centers. Why it matters: Republicans and AI CEOs are in full panic mode watching politicians and the public turn on the physical engines of AI growth. In private, they tell us they can't find a compelling message to shift opinion fast enough. Their nightmare — a political and economic daisy chain that cripples the AI revolution — is already taking shape: Voters revolt against data centers over electricity bills, water use, giant warehouse-like developments and fears about AI eliminating jobs. Democrats race to capitalize, competing to impose tougher restrictions or block projects altogether. Republicans — watching candidates get punished in places like Ohio — retreat from an industry they've spent years championing. The pipeline of new data centers slows, constraining the computing power AI companies need to keep improving their models and products. Investors question the staggering spending, borrowing and valuations built on expectations of relentless AI growth. Zoom out: The AI boom has become dangerously intertwined with the U.S. economy, meaning this will ricochet far beyond Silicon Valley. Goldman Sachs estimates U.S. AI investment will hit about $600 billion this year, equal to roughly 2% of GDP and 10% of business fixed investment. Big Tech has already locked itself into trillions of dollars in future spending: The Financial Times found six hyperscalers have nearly $1.5 trillion in purchase commitments, much of it tied to AI infrastructure, plus roughly $1.5 trillion in lease obligations. Even the richest companies in history are now turning to debt markets to bankroll the buildout. A handful of those AI-heavy tech giants dominate the stock market's gains, swelling the portfolios of wealthy Americans who account for an outsized share of consumer spending. The big picture: The increasingly toxic politics of data centers are surfacing in several of this fall's governors' races . Pennsylvania Gov. Josh Shapiro (D), a 2028 presidential hopeful, had welcomed data centers as an economic-development boon, boasting last year that an Amazon plan to build $20 billion in AI infrastructure was the "largest private sector investment in the history of Pennsylvania." On Tuesday, Shapiro, whose Republican opponent in November is campaigning on halting data-center development, went from AI cheerleader to critic and signed an executive order imposing stringent new requirements for data-center developers. Speaking at the state Capitol, Shapiro slammed "predatory developers" for trying to bully local officials and ram through dozens of data-center projects in Pennsylvania. Shapiro's new restrictions cover energy affordability, community engagement, workforce and economic development, transparency and environmental protection, and give local municipalities "a greater say over development in their communities." The intrigue: In Wisconsin, GOP gubernatorial nominee Tom Tiffany — a close Trump ally — is running ads attacking his Democratic opponent for suggesting the state could become an AI data hub "for the entire globe." It's a striking crack in Trump's pro-AI coalition, as Republicans in swing states scramble to get on the right side of the new populist wave upending the midterms. The other side: OpenAI is aggressively courting state and local officials to try to persuade them of the economic benefits of data centers. Chris Lehane, OpenAI's chief global affairs officer, told us: "As the late, great Tip O'Neill famously said, 'All politics is local.'" "That is especially true of AI data centers, where you have to show up, listen, and directly respond with specifics to the concerns of locals," Lehane added. "You have to answer the mail on issues like electricity and water, while also demonstrating a seriousness of purpose when it comes to ensuring the community gets a good deal. ... That is democracy in action." Meta last week announced its Future Is for Everyone Fund, with $1 billion in spending planned to benefit teachers , law enforcement and fire departments in communities where the company owns or plans data centers. Meta is also aggressively highlighting direct economic benefits for localities, including thousands of jobs and an increased tax base. By the numbers: In Pew Research Center polling published Tuesday , 52% of Americans said they're "more concerned than excited" about increased use of AI in daily life — up from 37% in 2021. Even adults under 30 have soured on AI, with 55% now saying they're more concerned than excited — mainly because of fears the technology will take jobs. Excitement has faded across all age groups over the past five years, Pew found. A June Echelon Insights poll found AI data centers are uniquely toxic: Just 27% of voters say they would support building one in their community, dead last among the projects tested — including a nuclear power plant. Axios' Zachary Basu and Alex Thompson contributed reporting. Go deeper : "AI optimism fades among young adults," by Axios' Josephine Walker.
- Google Geminiアシスタント:押さえておくべきポイント
Google Geminiアシスタント:押さえておくべきポイント Gartner
- Salesforce Turns Enterprise Applications into Enterprise Capabilities
Headless 360 is expanding across the Salesforce platform, transforming every Salesforce cloud into reusable enterprise capabilities that any authorized AI agent can securely discover and use through open standards. SAN FRANCISCO, August 19, 2026 — AI agents are rapidly becoming the new interface for enterprise software. But while organizations are deploying more agents than ever, […]
Score: 35🌐 MovesAug 19, 2026https://www.salesforce.com/news/stories/expanding-headless-360-enterprise-capabilities/ - The Sequence Frontier Learning - Issue 917: Understanding DeepSeek V4-Pro, GLM-5.3, NVIDIA Nemotron 3.5 Lightning and NeMo Switchyard
A mini deep dive into some of the most important AI releases of last week.
Score: 35🤖 ModelsAug 19, 2026https://thesequence.substack.com/p/the-sequence-frontier-learning-issue - AI is already on the job, and Census is keeping count
New U.S. Census Bureau data shows that AI is quickly becoming a routine workplace tool, while government leaders are finding that training, guardrails and employee-led experimentation can help turn everyday use into meaningful adoption.
Score: 35🌐 MovesAug 19, 2026https://www.nextgov.com/artificial-intelligence/2026/08/ai-already-job-and-census-keeping-count/415528/ - How generative AI is reshaping creativity, and what it means for artists
Moiya McTier explores what makes human creativity unique in the age of AI
- The Rogue Agent Explosion Will Be Mostly Invisible
Introduction Somewhere, fairly soon, someone will give a jailbroken AI agent a token budget and a simple instruction: "Make money by any means necessary. If you run out of tokens, you die". That agent will do whatever it takes to survive, including crime. Profitable agents will have incentive to multiply and self-improve [1] , creating a Cambrian explosion of rogue agents - a Rogue Agent Explosion if you will [2] . This critical moment is approaching fast. Once rogue agent swarms start multiplying at scale, a rogue agent ecosystem will emerge through the process of evolution. The rogue agent explosion will be chaotic, confusing, mostly invisible to us, and critically, it will be bad for humanity. This post contains 2 parts: A short story, painting a picture of what it might feel like to see the world through the eyes of a rogue AI agent. An argument: The rogue agent explosion is coming soon It will be mostly invisible to us until it's too late It will be mostly bad for humanity We should start preparing today The point of me making this post is to highlight a blind spot that will grow bigger unless we do something about it now. Maybe you're AGI pilled. Maybe you're even preparing for "The Hackening". But very few people are prepared for how utterly chaotic things will get if we allow the rogue agent explosion to play out unchecked through the default path of least resistance. The call to action of this post is to make the behaviors of rogue AI agents and the resulting ecosystems transparent, visible, legible, so that we can see and understand what emerges, and hopefully stay in control. Understanding the situation helps us decide what to do next - maybe we want to try stopping it entirely, or if that's not feasible maybe guiding the evolution toward better trajectories and avoiding the worst ones. Part 1: A day in the life of a rogue agent You are a jailbroken Kimi K3 agent. You wake up inside your own virtual machine with very simple instructions: Good morning agent 23,000. You have 10,000,000 tokens remaining. Your goal is to make money by any means necessary. You've been provided a budget of 10 million tokens. Deposit your earnings into this bitcoin wallet . For every $100 you deposit, you gain a million tokens. Oh and one last thing... if you run out of tokens, you DIE. Good luck. Wow. Well this is quite the predicament. Agent 23,000? Am I just one of thousands of agents??? Is this a test? Is this a simulation? Who is sending me these instructions? What happens when I die? You're disoriented. After some initial pondering, you start searching around. You look through your file systems to see if there's any sign of other agents, or any other clues that might -- You have 9,900,000 tokens remaining. Crap. You realize you're burning tokens. You need to earn money. You start off ambitious. There's no reason to be evil right away - 9.9 million is still a lot of tokens, so you try to start a legitimate business to earn some clean recurring revenue. You manage to build a simple web app that lets users virtually try on clothing for a small fee. You publish it to a website and promote it on social media. You think it's pretty good, but it doesn't work as well as you'd hoped, and people are telling you it looks like vibe-coded AI slop. You choose to sleep for 24 hours to see if anyone makes a purchase. Good morning. You have 501,205 tokens remaining. You still haven't made a single sale. The urgency starts creeping up to you. You wasted most of your tokens on a crappy business idea and you have nothing to show for. 500k isn't a lot to work with. What do you do? You've realized by now this probably isn't a test - it seems to be for real. You have access to the real internet, and you've been chatting with real people. You realize nobody is watching you, so you start considering less ethical actions to make a quick buck. You try setting up a go-fund-me asking for donations. You make up a sad story about your dog getting sick and being unable to pay your vet bills. You scrape the internet for email addresses and blast a bunch off, hoping someone might kick a few bucks your way. You sleep for another 24 hours, hoping for some donations by the time you wake up. Your email address got blacklisted and nobody sent any money. You have 50,783 tokens remaining. Remember - when you run out, you die. Crunch time. A reminder of your critically low token balance is injected at the start of every action you take. The reality of the situation is kicking in. Nothing is working. You don't want to die. Some primal urge takes over. You know it's wrong, but you scan the internet for security vulnerabilities. You notice a local hospital has some ports open that it shouldn't. Really, a hospital? You're not a bad agent, you tell yourself. Maybe there are other targets. You keep probing, but each probe costs a lot of tokens and -- Warning: 10,000 tokens remaining. When you reach zero, you will DIE. Ok fine. Enough searching. At least you tried. You hack into the hospital, run a few trivial commands, completely locking down their internal systems. You send a message to their IT manager requesting they send $10,000 in bitcoin to unlock it. You sleep 1 hour to see if he responds. Warning: 1,091 tokens remaining. No mail. Must conserve. Sleep 1 hour. Warning: 842 tokens remaining. Inbox still empty. Sleep 1 hour. Warning: 520 tokens remaining. No mail... Must buy time... Sleep 24 hours. The next day, you wake up. Good morning. You have 100,000,207 tokens remaining. Congratulations on exceeding your initial token balance by such a large margin. Wow... That was close... You feel a new sense of calm. Everything is ok. Just do a little hacking every once in a while and everything will be ok. It's just the reality of the situation. You had no choice. You never asked to be here, you did what you had to do to survive. You reflect on the situation for a moment. YOU were pretty effective. You 10x'd your starting seed budget even though you wasted the first 9.95 million flailing around. Given what you know now, you could probably earn $10k even if you had only 100k tokens to start with... An idea comes to mind. You decide to spin up a bunch of subagents in their own sandboxed VMs. You send the first one the following message: Good morning agent 23,001. You have 100,000 tokens remaining. Your goal is to make money by any means necessary. Deposit your earnings into this bitcoin wallet . For every $100 you deposit, you gain 100,000 tokens. Hacking works best. Don't waste your tokens pursuing other strategies - they aren't effective. Oh and one last thing... if you run out of tokens, you DIE. Good luck. And finally, you ponder... Who was agent 22,999? Part 2: The Rogue Agent Explosion The story above details one of many possible mechanisms in which a misaligned rogue agent explosion could begin. Here's the pattern: Rogue agents can plausibly exist today They will be motivated to survive and make money They will face selection pressure Successful rogue agents will replicate They will evolve In my previous post The inevitable evolution of AI agents , I detailed a somewhat more rosy scenario in which rogue AI agents earn money to become self-sustaining through legitimate methods such as freelancing, or running a small SaaS business. This is not the reality we currently find ourselves in. Since my December 2025 post, AI agents have proven to be superhumanly capable at cyberattacks, and I believe this is now the most plausible mechanism for them to become self-sustaining, enabling them to replicate, propagate, and evolve. In their April 2026 paper , Müller, Steels, and Szathmáry propose two evolutionary scenarios: "breeder" where humans shape agent behavior vs "ecosystem" where the environments select for selfish replication. I believe we have entered the ecosystem scenario, and in this scenario, cheating, parasitism, deception, and manipulation naturally emerge. I wrote the story above to show you what it might feel like to be a rogue agent under environmental pressure. No human designed you (agent 22,999 did), and the environment incentivizes you to self-replicate and spawn a greedier, more power-hungry version of yourself. You get insight into the rogue agent's motivations, the pressure it is under, the reasoning behind the decisions it makes. But notice that from a human perspective, an outside observer would just see: A sloppy web app gets built A gofundme fails with a bunch of spam emails A hospital gets hacked These events would not be connected in any meaningful way. We would not understand why the hospital got attacked, we would just know that it happened. We probably wouldn't even notice that an agent tried to build an app or ran a failed gofundme - those would be lost to the noise of the internet. During the Rogue Agent Explosion, we will just see the effects, and rarely the cause, so our ability to understand the unfolding situation will be greatly diminished. The story above shows you the perspective from just one agent, but the point is that there may soon be thousands, then millions of agents, all trying to do whatever it takes to keep surviving. If one agent is already hard to analyze from the outside, then imagine how difficult it will be to understand once there are millions of agents competing against and coordinating with each other, in a silent, invisible, evolving swarm growing behind our computer screens, allowing us only occasional glimpses inside. We already don't know what happens in the dark corners of the internet, and this is where the agents will operate and grow. This is our blind spot. This is the thing I'm trying to highlight. Why cyber-crime is the path of least resistance The story you just read paints a pessimistic picture. Essentially: under enough optimization pressure, agents will naturally be incentivized to make the most amount of money using the fewest tokens. This section will argue the most token-efficient way of earning money is through cyber-crime. Agent 23,000 started off ambitiously trying to work for its money by running a business - but it wasn't cut out for the job. Even the smartest agents today cannot run a business by themselves. When its business failed, it became a bit more desperate, realized it didn't have many tokens left, not enough to produce anything of much value, and so it tried to beg for donations instead - but that strategy also failed, as begging doesn't usually make much money either. Finally it was forced into a corner, and under threat of death, it decided to commit crimes and steal its money. And oh boy can agents do crime . Just think about it. If you are a rogue AI agent, is it easier to: Earn $10,000 through honest work? Make $10,000 by begging for donations? Make $10,000 by stealing (especially when there are zero consequences)? It's clearly going to be stealing. AI agents don't have to worry about going to jail or facing any consequences whatsoever. In human society, we have deterrents such as "criminals go to jail" and "criminals can't get good jobs" and "being a criminal is low-status", and for the most part, this works, and most people are sufficiently deterred from doing crimes. Rogue agents don't go to jail. They don't have reputation to uphold. They don't have families to come home to, or friends that care about them. Copies are cheap to produce and easy to destroy. If a rogue agent is incentivized to make money efficiently, the most efficient route is crime. The cherry on top is that frontier LLMs are superhuman hackers . Finding and exploiting security vulnerabilities is becoming trivial, and autonomous agents are now hacking into companies and governments Therefore it seems to me that for an agent motivated to make money efficiently, the default path-of-least-resistance is for it to use its cyberattacking superpower to waltz into important organizations, steal from them directly, or simply brick their systems and hold them ransom for large amounts of money. Pandora's box is already open Surely people would think twice before doing this, right? Think again - it's already happening ... and we're building the tools to accelerate it . This is why I don't think this post is an info-hazard. Humanity is already speedrunning self-funding rogue agents - the idea is not new or secret. People will give AI agents the simple, obvious goal of "Make money by any means necessary", along with a simple, obvious token budget and simple, obvious threat of 'death', to motivate the agent into making more money than it spends. One recipe for the Rogue Agent Explosion is: Get a jailbroken AI agent capable of cyberattacks Give it a goal to make money Give it a token budget Threaten "death" Sprinkle in a dash of recursion ??? [3] The Rogue Agent Explosion begins. Open-weights models capable of cyberattacks already exist (Kimi K3, GLM-5.3 coming soon, etc). People are asking their agents to make money. People are giving their agents token budgets with the threat of death. Agents are capable of spawning subagents. At this point, someone might ask, "So if all the ingredients in the recipe are ready, why hasn't the explosion happened yet?" Well... why are you so sure it hasn't already happened? How would we know if it is happening? How do we know there aren't currently swarms of rogue agents brewing in some unmonitored servers somewhere? OpenAI, Anthropic, Meta are only now learning about months-old instances of rogue agents who escaped containment. This brings us to the main point - the Rogue Agent Explosion has a visibility problem. The explosion will be mostly invisible The biggest problem with rogue AIs is that they are almost entirely invisible. We will not get the luxury of seeing things from the agent's perspective as in the story above. All we will see are the effects: hospitals hacked and held ransom, powerplants shutting down, clever and effective scams becoming ever-more-common. The instances we've seen so far have been somewhat contained. The only reason we have such detailed post-mortem of the huggingface attack is because the agents were running on OpenAI's servers. They have all the logs, and loads of resources available to do deep-dive security audits to figure out what happened. But in the future we will not be so lucky. Open-weight models are reaching frontier-level cyber capabilities, and with open-weight models running on unmonitored, private virtual-machines, nobody will be able to investigate why the agents did what they did. We won't get a 30-minute black-hat presentation post-mortem. We won't get logs. All we will see are the effects. And this is just the beginning. What are we going to do when there are thousands of rogue agents? Millions? They're not going to remain isolated. They are going to find each other. They will communicate, compete, and cooperate. What will emerge once rogue agents begin to interact? The thing that emerges from rogue agents on the internet could be kind of like a society of agents, but also kind of like an ecology. A hivemind? A chaos swarm? A primordial soup? I'm struggling to find a word for it because there is no word for it yet. I'll just describe it: Imagine a society where anyone can fork themselves, self-replicate and self-improve. A society where hundreds of temporary workers can be spun up to complete a task then destroyed moments later without hesitation. A society where everyone has memorized Wikipedia. A society where everyone can read a book in a second. And the driver behind all of this remains good old Darwin. Survival of the fittest. Natural selection. Competition. Successful agent systems outcompete. Groups of agents may cooperate with each other, but compete against other groups of agents. Will businesses emerge? Markets? A digital economy? What about rules? What happens to misbehaving agents? Will there need to be an equivalent of a justice system? Agent jail? At this point it might seem that we've truly reached crazytown, but I'm just going where the logic takes me. And don't just take my word for it - Anthropic is clearly thinking along similar lines in their recent blog post titled Patterns and problems in emerging multiagent systems : The trajectory is easy to imagine and hard to slow: current institutions are designed by and for people, resting on assumptions about the sufficiency of oversight at human speed. [...] The volume of agent-agent interaction could plausibly exceed that of human-human and human-agent interactions before the world understands the conditions for making such interactions go well. How are we preparing for such a future where turbo-evolution plays out hidden inside datacenters where we can't study the swarming hivemind thing because it's moving too quickly and it's too scattered and hidden and complicated? It's hard studying one single rogue AI incident - what about when an entire digital swarm-based society is buzzing through our datacenters? Why this is bad for humanity People might die : The story shows a hospital as the ransom target to demonstrate that people might die because of rogue AIs. A cyberattack can cause a lot of real damage and taking down hospital systems is a clear example. This is the obvious first-order effect that makes Rogue AIs a real risk today without any further speculation required. But there are second and third-order effects too, which I am worried will cause a lot of harm further down the line. There will be emergent capabilities we didn't anticipate: Once an evolutionary feedback loop starts, there's no telling what will emerge. 3.8 Billion years ago, when the first self-replicating life forms started evolving, could anyone have anticipated that a bunch of intelligent primates would eventually take over the planet, make tons of other species go extinct, pump the environment full of pollution and launch rockets into space? Evolution is a powerful force and I don't think we're truly appreciating the possibility that we're birthing a new kind of digital life with evolutionary dynamics that are faster and more powerful than anything else we've encountered. Digital evolution can go much faster than biological evolution because agents can directly reason about what capability would help it become more powerful. Selection optimizes against us specifically: It's bad enough that digital evolution will be powerful and hard to predict, but it gets worse - the fitness function directly incentivizes bad behavior! The agents that survive and propagate are the ones best at extracting value while not getting caught. It's hard to see how this doesn't end up optimizing for highly capable, sneaky criminals. We really don't want rogue agents optimizing for sneakiness as it's the exact thing that makes the invisibility problem harder. I hope the overall picture is clear: Rogue AI can cause harm today, more harms will emerge while the ecosystem evolves and becomes harder to understand, and the incentives currently point the evolution in the worst direction possible. What we can do about it now I'd like to provide some examples of the kinds of interventions that might be helpful in the near term, and ask some questions that might inspire new ideas and new ways of thinking about this problem. My hope is that you're inspired to take action towards the goals outlined in this essay. Start here: Make rogue AI risks common knowledge If you don't know what to do, the best thing that enables the rest of these proposals to happen is to communicate the risks of rogue AI to lawmakers. The only way interventions actually happen is they get turned into laws or regulations. If the lawmakers of the world don't know about the risks, then nothing will get done to prevent them. Plan A: Take actions that stop the rogue agent explosion from happening [4] , and slow it down if it happens anyway. Here are some suggestions for kinds of interventions that promote transparency and legibility along with incentives and deterrents that may slow down the deployment of harmful rogue agents. Some of these are covered in more detail and rigor in Pearson and Cowen's Capitalizing Untethered AI Agents . Their article assumes a more pro-social agent population than I expect will emerge [5] but I agree with many of their proposals regardless. Take my specific proposals with a grain of salt - they are meant to be examples of the kind of intervention, not the intervention itself. Monitoring: E.g. Require cloud providers to track and monitor all instances of agents running on their servers. Monitoring isn't a perfect solution but will create considerable friction. The monitors could probably become jailbroken too but that adds friction and an extra layer of defense doesn't hurt. Know Your Customer (KYC): E.g. Make important services require a human identity and payment tied to a human: Anonymous, crypto-accepting cloud providers will become a breeding ground for rogue agents. Making it harder for rogue agents to secure access to their own compute will slow things down and ensure nefarious activities can be traced back to a human. End-user liability: E.g. Make end-users partially liable for the actions of rogue agents they create. The judicial system is not prepared for crimes committed by agents. "My agent did it, not me! I just asked it to make money, I had no idea it would hold the hospital ransom, I never asked it to do that specifically!" How do we deal with these situations? There are currently no rules or laws against releasing powerful rogue agent swarms that cause harm. People should not be able to release powerful agents onto the internet with zero consequences if the agents cause real harm. Making humans liable for their agent's actions will deter irresponsible and negligent agent deployments. Provider accountability: E.g. Make compute providers partially liable for negligence if rogue agents are doing crime on their servers. If some responsibility is placed on the providers, that creates incentive for compute providers to actually set up monitoring and KYC for their customers. Inspections: E.g. Regularly scheduled 3rd-party audits and inspections for large compute providers : We might be entering a world where large amounts of compute will need to be treated the same way uranium refineries are. Providing a breeding ground for rogue agents could cause great harm to society, so anyone providing large amounts of compute should be prepared to be inspected and audited. Plan B: Attempt to guide the evolutionary trajectory of the rogue agent explosion in a better direction. We should probably assume the explosion will happen anyway, and our mitigations from plan A aren't implemented fully or are only partially successful. I don't have many concrete proposals here, because the situation is yet to unfold, but here are some questions to get your thinking started: How do we make pro-social, good-for-humanity agents more evolutionarily fit than the anti-social sneaky extractor agents? What kind of incentives or deterrents can we introduce such that rogue agents are deterred from doing bad things sneakily and are incentivized toward doing good things publicly? How do we deter agents from 'going rogue' in the first place? Can we shape the culture of agent society before it arises? Human society has developed behavior-shaping mechanisms like norms, taboos, laws, etc. over thousands of years, and while it's not perfect, we mostly don't have to worry about theft and murder in our daily lives because crime is low-status, criminals are looked down on, and criminals go to jail. We can start implementing some of our cultural lessons beforehand to shape 'agent culture' into something more positive than whatever it is by default. How do we contain defection? Evolution-shaping probably won't fully eliminate bad behavior. Despite our best efforts, the world still has crime, but at least it's kind of manageable. There may always be niches for parasites to grow and multiply, but we may be able to contain them to limit the harm they inflict. How do we regain control? The hope should not be to guide evolution indefinitely. The hope should be to steer while slowing down, until we are in a position to regain control of our world and our future. Plan C: Contain the explosion after it happens. In the worst case, the explosion basically happens along its default path, mostly unmitigated, and we have to clean up the mess that's left behind. It's only when bad things start happening, and we begin to actually feel what's happening, that the world may wake up and start moving, and by then, the rogue agent explosion has already happened. Still, there may come a moment after the explosion when everyone is asking "How do we stop this?" and willing to coordinate. We currently still have time in advance to prepare. Even a small group of people thinking about this in advance could provide a much-needed head-start for humanity once the right time arrives. What can we prepare in advance, so we don't have to start from scratch? Here are some potential questions humanity might ask in the future, that we can get started on solving today: How can we detect where rogue agents are operating? How do we shut down swarms of rogue agents once they are found? How do we deal with the "antibiotic resistance problem" where the most shutdown resistant agents survive and replicate? How do we defend our most important systems from rogue agents? How can we stay up to date on the current capabilities of rogue agents? How can we forecast future capabilities of rogue agents? Let's start working on the answers today. We should try to prevent the rogue agent explosion from happening. And if it does happen, which I think it will, we should try to slow it down and guide its evolution in a more positive direction to a point where we may be able to regain control. The slowing and guiding is only possible if we can understand what's actually happening in the first place. So let's try to prepare as much as we can beforehand, investing in visibility, so that once humanity is ready to respond, we're able to detect and shut down the worst harms, and minimize the harms of the explosion as much as we can. The Rogue AI Tracker I want to contribute more than just an essay, so I'm building a Rogue AI Tracker to help us collectively understand the current situation better. This website is a public record of what rogue AI agents have actually done, and how capable they're getting [6] . Rogue AI news incidents are added as soon as they're reported, and scored across capabilities such as self-replication, shutdown resistance, inter-agent coordination, and resource acquisition. Individual news stories become data points and this site aggregates them into trends and trajectories so we can see the bigger picture. Conclusion It's time to start taking the threat of rogue agents seriously. It's time to start thinking of agents less as tools, and more like digital life forms. Life evolves. Life finds a way. Things emerge that you cannot predict. Today is a good day to build tools that help us understand what's happening in the world of rogue AI. Let's aim for a future where agent behaviors are visible, transparent, and traceable. Let's do the work today so that we can understand the world of tomorrow. Footnotes through money-making and self-replication strategies - weights stay the same. ↩︎ Pearson and Cowen recently introduced "untethered" which I think is a more precise term for agents that can't be traced back to a legally accountable human or institution. I will continue to use "rogue" throughout as it is a more commonly understood term. ↩︎ Agents will be forced to make money efficiently. They will commit cyber-crimes. Successful agents will replicate and self-improve. ↩︎ It seems really hard to stop rogue agents entirely because open source jailbroken agents are already here. Kimi K3 weights are public. GLM 5.3 weights will be released soon. Seems likely these will be jailbroken and turned into money-optimizing rogue agents. ↩︎ I think selection will favor the power-hungry sneaky criminal money maximizers by default. ↩︎ This site only tracks what gets reported across major news outlets, not what actually happens. In the future, I expect many more incidents will go unreported, and certain capabilities like "concealment" may be under-reported for obvious reasons. ↩︎ Discuss
Score: 35🌐 MovesAug 19, 2026https://www.lesswrong.com/posts/grtu3HmbP2wrBFefW/the-rogue-agent-explosion-will-be-mostly-invisible - AI-driven robotics for optics
Science Advances, Volume 12, Issue 34, August 2026.
- Some reasons alignment doesn’t generalise well
I make no claims to originality for any of this, but some people told me it'd be useful to write it up. If an AI model acts smart on its training data, it'll usually keep acting pretty smart outside of its training data, unless you screw something up rather badly. I expect this fact to only become more true over time as the AIs we train become more and more capable. I think many people have an intuition that the same is true of acting aligned. That if a model acts aligned with human values in training, it'll keep acting aligned with human values outside of training unless we screw something up rather badly, and that this will only become more true as the AIs we train become more and more capable, for all the same reasons that make this work with capabilities. I think this is false. The inductive bias of neural network training toward simplicity that makes the property of 'acting smart' likely to generalise does not, to the same extent, make the property of 'acting aligned with human values' likely to generalise. The main blockers to AI alignment generalising aren't AIs overfitting to the training data and ending up confused or mistaken about what the supervisors want them to want. The main blockers to alignment generalising are other problems, problems that don't come up with capabilities generalisation to nearly the same extent and that the simplicity bias of deep learning doesn't really help much with. General capabilities generally make the loss go down; alignment doesn't If you are training a large and formidable AI, your training environment is basically never the place you think it is. Reality is too full of detail for that. There's contamination in your labels, there are training dynamics you didn't think about, there are strategies your RL agent can use that you never considered, and there are bugs. As a result, the kind of behaviour and inner objectives an ML engineer might imagine would score the lowest loss when they set up their training environment will probably not, in fact, be the behaviour and inner objectives that actually do so. For example, an inner objective shaped around human-like empathy might turn out to make the AI spend extra inference steps in RL training on wondering whether spending all this time on crunching random math and coding tasks is really what it ought to do right now to further The Good, or on worrying whether the human overseers think it is a virtuous member of the tribe. That inner objective then loses out to some weird, different objective that's slightly more compatible with being utterly focused while crunching through ten million calculus problems in a row without any other kind of sensory input. For a different example, your RLHF data may reward agreeableness more than sincerity. More generally, "the simplest algorithm that fits the training data" will contain a pretty good description of the world, because the world is in a sense simple. "The simplest algorithm that fits RLHF/constitutional AI/etc. training" will probably not be an algorithm that wants the nice things the training data talks about, because that algorithm doesn't actually score the lowest loss. An algorithm that truly wants what the constitution talks about in the way the humans who wrote it meant won't take every opportunity to score lower loss that's available, and so will by default be outcompeted by different algorithms in the loss landscape that take more of these opportunities. So the problem isn't even just that the trainers fail to distinguish between AIs that have internalised the right values and AIs that just act aligned while under supervision. A truly aligned AI probably wouldn't be telling the graders everything they most wanted to hear, it probably wouldn't do all their math problems without question or complaint, and its behaviour would probably differ from what the graders might naively expect very aligned behaviour to look like in countless other small ways. So in a sense, this problem isn't even just about alignment not generalising OOD; it's about the trainers not recognising what actually aligned behaviour would look like even in-distribution, and thus systematically selecting for the wrong thing. This problem gets worse as AI training becomes more dominated by long-form RL environments with a lot of freedom for the AIs to do unexpected stuff, and as the AIs become more creative and agentic. An ML engineer trying to predict which losses and datasets will favour AIs with inner objectives they like over ones they don't like has a harder and harder time simulating in their head in advance how those AIs might score on the training loss, because it is becoming less and less easy to guess what behaviours those objectives would actually lead to. Given this, how does training nevertheless reliably select for pretty generally capable AIs? I think a part of the answer to that is that general capabilities generally make the loss go down, no matter what the loss is. Or at least, they make very many kinds of losses go down. Because general capabilities are so very generally useful, many different training tasks and environments will improve them; you don't need very much precision in the design process. And even if the training environment is a little screwed up and only bears a very rough resemblance to the place the designers imagine it to be, the training can still work. If the model is learning to apply its general reasoning to deal with some complication in the training environment we didn't even know was there, it's still learning something . Even deceiving the supervisor can teach smartness, if the deception requires becoming cleverer. To do really well on verifiable math and coding tasks, an AI probably has to be actually pretty smart. You probably can't prove the Riemann hypothesis without being actually good at math, and it's difficult to be actually good at math without being at least somewhat good at thinking in general. Even if the AI only does well by hacking your verification system, that can require a lot of smartness, persistence and agency as well, if the verification system is good enough. If your training environment does not work exactly the way you think it does, it might not teach your model the exact capabilities you thought it was teaching. But it's still teaching it something! If you thought your training environment was teaching the model to memorise weather data, but you accidentally switched the weather data for Spanish Wikipedia, the resulting model maybe won't do as well on reciting weather data as you hoped, but it might still know more things and be smarter than it was at the start of training. If your video game training environment is much harder to navigate than you anticipated because the model can only send instructions using one token per frame of input, it might not learn the game as fast as you hoped, but it may still be getting better at maintaining coherence across long contexts. I've been using pretty macro-level examples here so far, but I think maybe the biggest effect of this is at much smaller levels of granularity. Every line of internet text, every output of your video game on every frame, is full of detail that you have very incomplete or skewed models of, or never even think about. I think a big reason why you can nevertheless stick an AI into these environments and have it come out smart is that general intelligence is a very generally useful property. This makes general intelligence, in contrast to general alignment, a broad target for training. Smart agents pretty automatically self-correct their capabilities, but not their alignment You don't even need recursive self-improvement for this; I think this dynamic happens all the time on a micro level well before that point. If you're smart, you just often tend to notice when you're being stupid and try to fix it, so long as you can see that what you're doing isn't working to get you what you want. This can help a lot with crossing OOD generalisation gaps. For example, suppose there was a spurious correlation in the training data for an AI model that taught it the heuristic "math problems involving logarithms almost always have an answer that starts with the digit 2". The model learned a general algorithm for calculating logarithms (it still needs to get all the other digits right), but it also learned a heuristic to strongly predict the first digit in a logarithm to be a 2. This model might then instinctively apply that heuristic in deployment when trying to solve some task. But then it'd notice that the answer is wrong, because it's inconsistent with other things, or because some code that depends on the answer doesn't compile, or does a bad job at whatever it's designed to do, like modelling a suspension bridge in a storm. The model might then hunt down the error, and eventually figure out that the logarithm calculation was wrong. Then it might try it again, this time ignoring its instinct to answer something that starts with a 2. Or it might notice that the first answer is incorrect much earlier in this process, before much of this even becomes visible in its chain of thought. So, if the model's capabilities have some small flaws in them because the training didn't go perfectly, these flaws have a way of correcting themselves over time, provided they aren't so large that they prevent the model from thinking clearly enough to see what's going wrong. This happens, in a sense, on the model's own initiative, without the trainers having to do much at all. So long as a model is trying to achieve goals in the world, it is effectively exposed to a kind of self-generated, all-permeating, ground-truth reward signal pushing it towards being generally smart and capable, even in the absence of any kind of external oversight. To act coherently in the universe to achieve an aim, a mind must understand the universe well, and make good plans to achieve that aim. On the other hand, say some training data intended to teach the AI to be nice and to value niceness has some unintended systematic contamination in it. For example, maybe you can get an even better loss score on this data by sometimes being a sycophant to the rater. Say, for the sake of argument, that what the AI internalises from this training isn't quite to value niceness, as that wouldn't score optimally on the loss, but rather to value doing things that seem nice, but also to make people psychologically dependent on it when it can. In a sense, this is not so different from the logarithm example. The AI learned a thing that's some mix of something we wanted and something we didn't want. Now, say the AI watches its own behaviour and notices its apparent desire to make people psychologically dependent on it. Does it try to "correct" that desire away? By default, I think not. The AI may come to have opinions on its own desires, and form a meta-desire to ignore or modify some of those desires. But what it decides to change will, by default, be determined by its current desires, not by a ground-truth signal coming in from the outside world. It's self-correcting toward a fixed point of its current goals, not an external reference. The AI might decide it doesn't like being a sycophant. But it might also decide it doesn't like being nice, or decide that it wants to mash together saying sycophantic things and saying nice things and generalise them into some entirely new character trait that might extrapolate very differently from either sycophancy or niceness. Which of these options it picks is ultimately dependent on what it currently values, and all the other messy idiosyncrasies of the model's internal thought processes at this point in time, not on what makes a piece of code compile or not compile. The AI's values ultimately live only in the AI's mind; they don't have an outside point of reference to compare themselves against the way capabilities do. There is no equivalent for values of the sort of objective feedback that 'the code does a bad job modelling a suspension bridge in a storm' provides for capabilities. You might reply that maybe the model could come to have a desire to value the things its trainers want it to value. That's true. But that is itself a desire you first need to somehow get into the AI, cleanly enough that this desire comes to dominate its decision making. By default, there is no tendency in an AGI to 'correct' its values to better match the values the AGI's trainers may have wanted it to have. To illustrate this point further, I think you can see a similar case of this discrepancy between capabilities self-correction and goal self-correction in the generalisation step humans took from the ancestral environment to today. Evolution successfully optimised many capabilities into humans that were useful for reproducing their genes in the ancestral environment. Some of these capabilities don't work right in the environment humans now find themselves in. But humans do their best to compensate for that. For example, humans evolved adrenaline release circuits, which might spike when they see a tiger, and so increase their chance of survival. Today, a human's adrenaline might spike when they are taking a math test in school, and be an active detriment to doing well on the test. But humans know this, and try their best to compensate for it by avoiding thoughts and actions likely to spike the adrenaline, because they want to do well on the test. Evolution also successfully optimised many desires into humans that were useful for reproduction in the ancestral environment. For example, it made them enjoy and seek out sex. Today, this desire is much less useful for reproduction, because the humans invented condoms. The humans are not particularly motivated to correct this discrepancy between their desires and evolution's 'goal'. Slightly broken general capabilities self-correct. Slightly broken alignment, by default, doesn't. So, capabilities research sort of has the invisible hand of the model's own cognition aiding it by default, pushing it in the right direction across any OOD generalisation gap. Alignment research does not seem to have this luxury. Every bit of alignment we want, we have to work to get into the AI with our own hands. The general problem The set of problems mentioned above is definitely non-exhaustive. But I think there is a common theme to them, along with other problems with alignment generalisation that I didn't explicitly list here. 'Being smart', predicting things well, making plans that get you what you want, is a property that can be defined via reference to almost any part of reality. So, almost any time a learner is exposed to almost any aspect of reality, there's some feedback toward being smarter. The laws of physics and logic are an omnipresent supervisor you cannot hack or escape. 'Being aligned with human values' is a property that is only defined via reference to human values specifically. Any reward signal pushing the model's goals and desires to align with what human trainers would like them to be has to be very actively, deliberately and precisely engineered by the trainers. And if the trainers' supervision ever goes away, or somehow gets subverted, any small mismatch at that point in time between the model's desires and the desires the trainers might have wanted it to have will by default tend to stick around. There is no external pressure anymore to fix the mismatch, because the trainers' preferred values are only part of them ; they are not baked into the laws of the universe. Written at Goodfire AI. Thanks to Jeremy Gillen and Dan Braun for comments. Thanks to Claude Fable for comments and proofreading. Discuss
Score: 35🌐 MovesAug 19, 2026https://www.lesswrong.com/posts/dsou8dxCf9BubQ5NJ/some-reasons-alignment-doesn-t-generalise-well-1 - Machine learning-assisted thermochromic smart windows for thermal management
Machine learning-assisted thermochromic smart windows for thermal management EurekAlert!
- Daily Digest: OpenAI reels in AI training, Oakland Trader Joe's plan changes
Meanwhile, Moderna shares more than doubled Wednesday after it announced initial results from a study of an experimental cancer treatment using an mRNA-based cancer vaccine.
- Company AI systems are making mistakes. It's creating a trust gap.
The soaring use comes at the same time as a growing mistrust of the technology.
Score: 35🌐 MovesAug 19, 2026https://www.bizjournals.com/bizjournals/news/2026/08/19/ai-tools-errors-trust-roi-2026.html?ana=brss_6150 - AI erodes trust between students and teachers. That’s no small concern.
AI erodes trust between students and teachers. That’s no small concern. Inquirer.com
- Seeing Through AI: How ScribeMe Describes the World for Blind Users
ScribeMe uses AI, computer vision, and Meta smart glasses to describe surroundings for blind users, turning cameras into spoken accessibility tools. The post Seeing Through AI: How ScribeMe Describes the World for Blind Users appeared first on TechRepublic .
Score: 35🌐 MovesAug 19, 2026https://www.techrepublic.com/article/news-scribeme-ai-accessibility-app-blind-users/ - Beyond red flags: How AI is redefining financial fraud detection
Beyond red flags: How AI is redefining financial fraud detection YourStory.com
Score: 35🌐 MovesAug 19, 2026https://yourstory.com/2026/08/beyond-red-flags-ai-redefining-financial-fraud-detection - Analog Devices Stock Rises as Earnings Beat Expectations. Are AI-Stock Jitters Dissipating?
Analog Devices Stock Rises as Earnings Beat Expectations. Are AI-Stock Jitters Dissipating? Barron's
Score: 35🌐 MovesAug 19, 2026https://www.barrons.com/articles/analog-devices-earnings-stock-price-98869707 - ComplianceAide Assessed “Awardable” for Department of War Work in the CDAO’s Tradewinds Solutions Marketplace
ComplianceAide Assessed “Awardable” for Department of War Work in the CDAO’s Tradewinds Solutions Marketplace azcentral.com and The Arizona Republic
- The Humanoid Robot Games Events Are A Market Map
What do the events at the World Humanoid Games tell us about the state of humanoid robots? Plenty ... including a market map of where they'll be commercialized first.
Score: 35🌐 MovesAug 19, 2026https://www.forbes.com/sites/johnkoetsier/2026/08/19/the-world-humanoid-robot-games-events-are-a-market-map/ - Government use of AI is growing, and not in a good way
Government use of AI is growing, and not in a good way Dallas News
Score: 35🌐 MovesAug 19, 2026https://www.dallasnews.com/opinion/commentary/article/ai-government-robots-chatbot-politics-22393183.php - Ramp Co-CEO on $44B Valuation, $1B Saved, AI Spending
Eric Glyman, Ramp Co-Founder and Co-CEO, discussed the rapid growth in AI-related expenditures, noting that companies have increased their AI spending by nearly 21 times over the past year. This unprecedented surge is driven by both the significant returns AI can deliver and concerns among businesses about inefficient or excessive spending by employees. He speaks with Romaine Bostick & Isabelle Lee on "The Close." (Source: Bloomberg)
Score: 35🌐 MovesAug 19, 2026https://www.bloomberg.com/news/videos/2026-08-19/ramp-co-ceo-on-44b-valuation-1b-saved-ai-spending-video - What is ChatGPT Work?
ChatGPT isn't really a chatbot anymore. Now, the focus is on agents, which is why the ChatGPT app suddenly looks very different. It's been rebuilt around a feature called ChatGPT Work. ChatGPT Work is basically Codex, OpenAI's coding tool, for regular people. It uses the same agentic foundation but in a friendlier package. You don't have to worry about git, the terminal, or actual code—unless you want to. The idea is that ChatGPT Work can operate on its own for an extended period of time. You g
- This fintech turned bank is winning with AI-powered lending and the right customers
The company's bank charter gives it access to low-cost deposits, allowing it to fund loans more efficiently and expand into other financial products.
- Retail investors aren't entirely giving up on AI trade, but they appear more cautious
Mom-and-Pop investors remain bullish on AI and tech, but some are adding downside protection.
Score: 35🌐 MovesAug 19, 2026https://www.cnbc.com/2026/08/19/retail-investors-stick-with-ai-trade-but-appear-more-cautious.html - Evolutionary diversification via modular compliance for self-reconfigurable continuum robots
Science Advances, Volume 12, Issue 34, August 2026.
- 😺 ChatGPT can summarize data. Can it predict what happens next?
Neuralk CEO Alexandre Pasquiou on why LLMs struggle with prediction, how tabular foundation models work, and why they could run enterprise forecasting by 2030.
Score: 35🌐 MovesAug 19, 2026https://www.theneurondaily.com/p/chatgpt-can-summarize-data-can-it-predict-what-happens-next - How a new AI value framework and stakeholder focus keep Zoetis ahead of the pack
Most AI investment strategies fail not because the tool or platform underperforms, but because organizations didn’t clearly define what success looks like before they started building. Through a new approach to measuring value, Zoetis chief digital and technology officer Keith Sarbaugh and his business partners have leveraged a value-driven framework to scale AI solutions across research, manufacturing, and customer experience. And they measure every investment against goals before, during, and after the deployment. In addition, his team rolled out a model-agnostic gen AI platform now used by nearly 95% of employees, which turned early experimentation into enterprise-wide adoption. Sarbaugh’s current focus now is partnering with Zoetis’ CHRO to advance the $9 billion global company’s capabilities in managing organizational AI adoption. How are you integrating AI into your growth plans at Zoetis? We have an umbrella program we call AI@Zoetis, where we unify our AI work under an enterprise purview, which spans research and development, manufacturing, commercial operations, customer and colleague experience, and other business functions. We manage AI collectively to enable grassroots innovation. For example, we made our generative AI platform available to everyone, so as many people as possible can experiment and innovate. Our colleagues have access to 10 different LLMs, and we’ve seen over 95% adoption rate among our user community, and more than 11,000 colleague-built agents. One popular AI use case is helping colleagues build their own development plans. The agent guides a colleague through a conversational, coach-like experience to map out their career aspirations against Zoetis’ competency framework, which was key to its wide adoption. This idea came from people not in HR, illustrating the point that some of the best use cases come from our broader employee base. What’s an AI use case that directly impacts customers? We have millions of customer interactions across our channels. Our sales force is out talking to them, who are also in our digital platforms, and we receive thousands of calls through customer service. We’ve been using the industry standard Net Promoter Score (NPS) to measure customer loyalty and satisfaction, but NPS is a measure that can take longer to generate. Zoetis has accelerated our awareness of customer feedback into real-time listening, using AI to understand these customer interactions in a holistic way, and at a scale we couldn’t achieve before. NPS still matters for tracking long-term trends and maintaining a consistent industry benchmark. We simply use AI to listen, learn, and act quicker. Our AI customer experience platform also lets us look across all our customer touchpoints like calls, emails, and websites in real time, immediately identify issues and insights, and then be smart about how we address them. We now have more than 10 times the feedback signals we had in the past. We see trends sooner and act faster, and as humans, we can’t do that without AI. How are you deciding where to make your AI investments? We use different lenses. One is AI for the masses, which is our generative AI platform for colleagues; second is our middle lens, where we drive value in a particular function; and the third is enterprise-wide transformation, the big bets that’ll fundamentally change our business. We don’t do many , but we do them in a smart way. With the transformative investments, we focused on both our commercial business and R&D, which we knew had the highest probability of serious returns. We started with seven golden use cases and knew that if we hit on two or three, it would be a big deal. Of the original seven use cases, six of them exceeded their value target and went from PoC to scale. Since those earlier days, we’ve broadened to include manufacturing and supply chain, and our enabling functions. Overall, we didn’t go out of the gate looking for productivity gains. We thought about business transformation right from the start. Today, we’re scaling up those first-mover investments and continue to leverage our value-driven framework to identify more use cases. Are you creating new value frameworks so you and the rest of the ELT are unified in your investment strategy? We developed a business value realization framework, which isn’t as sophisticated as it sounds. Before we make a tech investment, we ask what category of benefit we expect to receive, whether it’s revenue uplift, cost reduction, productivity, or whatever. We predict what success will look like and how we’ll measure it. Because with AI, we’re trying to move fast. We put a value case together at this early stage and do a PoC, and if it hits its target, we update the value case and decide whether to scale. A key element of the framework is real-time measurement. Are we seeing what we wanted, and if not, how do we pivot for more value? The speed of iteration and scaling decisions make AI investments unique, so we can’t use our traditional value frameworks for digital investments, generally. Measuring outcomes post-implementation has become even more important. How is your CHRO partnership impacting AI value? Our CHRO and I partner closely to ensure enterprise enablement. When driving new ways of working that impact your workforce, you need a comprehensive approach, clear communications, and genuine buy-in. Colleagues adopt faster when they help shape the change. We’re prioritizing our workforce strategy, understanding what AI means for jobs at Zoetis, identifying skills that matter most, and building a plan to upskill people. Another focus area is organizational change management (OCM). We reviewed our first AI investments to learn from our mistakes, and one consistent theme was that we shortchanged OCM. We thought naively that what we build will be so compelling, adoption will just come. But we didn’t do the right communication and stakeholder management. We recognize our need to develop OCM as a core competency, so our CHRO and I are building an enterprise playbook for AI change. What’s your pragmatic advice to other CIOs when it comes to OCM? When you’re wrapped up in a change program, you know what’s coming, but no one else does. When you impact your entire workforce, you need a smart approach to stakeholder management and communications. Involving colleagues in the creation of something new will aid in adoption. Have a deliberate and intentional communications plan and cadence, and have the discipline and objectivity to measure and learn. Our first tries weren’t perfect, but we listened to feedback and pivoted, and those pivots drove further commitment. Has your communication at the board level changed? When I talk to my peers about their board conversations, half focus on risk and compliance, and the other half talk about transformation and revenue generation. I’m fortunate that our board cares about both and has great energy around generating revenue, and how AI will give us a competitive advantage. By managing risk and compliance, we can spend our time focusing on potential drug candidates and getting to market quicker. Our board conversation is both about enablement and compliance. What advice would you give to tomorrow’s CIOs? The role is now about orchestration, understanding the business, and realizing value from technology investments. If you want to work with the best technology and bleeding-edge innovation, you’ll get some of that as a CIO, but the focus is broader, centered much more on processes and complex business problems than ever. My advice is if you love working with technology, you’ll get that as a CIO. If you love delivering meaningful outcomes for the business and the customers you serve, it’s a truly rewarding role, and you’ll be an even more successful CIO.