AI News Archive: September 2, 2026 — Part 7
Sourced from 500+ daily AI sources, scored by relevance.
- OpenLeash Adds a Human Check to Risky AI Agent Actions
The security tool intercepts potentially dangerous agent actions, blocking clear threats and requesting human approval when intent is uncertain. The post OpenLeash Adds a Human Check to Risky AI Agent Actions appeared first on SecurityWeek .
Score: 31🌐 MovesSep 2, 2026https://www.securityweek.com/openleash-adds-a-human-check-to-risky-ai-agent-actions/ - Unitree employees reportedly describe founder-led management and high internal pressure
Current and former Unitree employees have reportedly described a tightly centralized management style and a demanding work environment at the Chinese robotics company. Accounts circulating online alleged that expenses above RMB100 required approval from CEO Wang Xingxing and that he was closely involved in product details. The reports also described strict confidentiality requirements and demanding […]
- Indian Startup HrdWyr Builds AI-Native SoCs for the Physical World
HrdWyr is developing AI-native SoCs for power management, motor control, and other applications where AI meets physical systems. The post Indian Startup HrdWyr Builds AI-Native SoCs for the Physical World appeared first on EE Times .
Score: 31🌐 MovesSep 2, 2026https://www.eetimes.com/indian-startup-hrdwyr-builds-ai-native-socs-for-the-physical-world/ - 🎥 Orchard Robotics CEO: Growers ‘shouldn’t have to be data analysts’
“In the first year, if a grower sees full value from labor and input savings, they can see anywhere from three to 10x ROI," says CEO Charlie Wu. The post 🎥 Orchard Robotics CEO: Growers ‘shouldn’t have to be data analysts’ appeared first on AgFunderNews .
Score: 31🌐 MovesSep 2, 2026https://agfundernews.com/%f0%9f%8e%a5-orchard-robotics-ceo-growers-shouldnt-have-to-be-data-analysts - Operational reliance on AI in supply chains and emerging insurance risks
Artificial intelligence (AI) is becoming embedded in supply chains. This report examines how that changes operational reliance, resilience, and insurance risk across logistics, transport, and warehousing.
- How the Kansas Department of Labor has implemented AI to make it easier to file for unemployment
Until a couple of years ago, the software running the state unemployment insurance platform in Kansas hadn’t been updated since the 1970s. Customer service agents had to keep five screens open at once while trying to help people filing for unemployment compensation; one of the tools was so old that it wasn’t even possible to use a mouse, and agents had to type commands into a command line. At night, the whole system had to shut down for batch processing, so if someone who’d lost their job tried to submit a claim at 9 p.m., they’d be frustrated to find that it wouldn’t work. The system finally got a revamp at the end of 2024—and now it’s continuously improving because of the implementation of artificial intelligence tools. Internal chatbots help customer service agents answer complicated questions from claimants or employers more quickly. “By design, unemployment insurance is extremely complex, and every claim is like a fingerprint—it’s unique,” says Amber Shultz, Kansas secretary of Labor, who has been leading the Department of Labor’s technological evolution. The AI tool searches internal information, forms, and other resources in real time while an agent is on the phone. To build the system, the team also used AI to search through its vast archive of documents, cleaning up duplicative or outdated information. As it’s used, the tool also helps the agency find commonly asked questions that it can answer on its website. Newly implemented chatbots also help people who are unemployed navigate the service, and as a result more than 90% of claims now don’t require talking with an agent. For those that do, staff has time to answer every phone call. In all, the changes have reduced benefits processing times by nearly 80%. The team is working on predictive AI tools, Schultz says, that can anticipate high call volumes and suggest that the call center temporarily add more representatives. The agency uses AI to create dashboards to help managers monitor and improve performance. AI tools also help flag fraudulent applications, and can find software vulnerabilities to protect the system from attacks by hackers. The agency is quickly finding new uses, though it’s intentional about the rollout. “We’re not just chasing the technology,” says Shultz. “We really look and evaluate to see if it can improve the experience for claimants and for staff.” It’s not likely to replace state employees, she says, but just to make it possible for a lean team to accomplish more with the existing workforce. The overall system will continue to change. “I think the thing that I’m most proud of is that we really changed the mindset of what government technology can be,” she says. “We aren’t just replacing those old legacy systems anymore. We really are building something that is more agile and data driven. And we as an agency are more willing to embrace any technology because we’ve seen what it can do for us.”
- Askya AI Growth Platform offers $200k zero-equity funding to African startups
Askya Investment Partners, an African venture capital firm, has announced the launch of the Askya AI Growth Platform, a pan-African programme to scale artificial intelligence startups into enduring, continental-scale enterprises. Askya Investment Partners is an investment firm building a new generation of technology and artificial intelligence champions from Africa, led by founder and managing partner [...] The post Askya AI Growth Platform offers $200k zero-equity funding to African startups appeared first on Disrupt Africa .
- How small AI models can make a big impact for enterprises
Explores how lightweight AI models benefit enterprise use cases.
Score: 31🌐 MovesSep 2, 2026https://cohere.com/blog/how-small-models-can-make-a-big-impact-for-enterprises - Asus puts a petaflop of AI ambition into ProArt laptops
Asus is taking its ProArt lineup deeper into the AI era with the new ProArt P16 and ProArt P14, the first ProArt laptops powered by Nvidia’s RTX Spark platform. The company is showcasing the machines at IFA 2026 in Berlin, positioning them as portable workstations for creators and developers who increasingly want to run AI […]
Score: 31🌐 MovesSep 2, 2026https://www.digitaltrends.com/computing/asus-puts-a-petaflop-of-ai-ambition-into-proart-laptops/ - Anthropic Has Some Alignment Problems
Oh, good. They noticed . Anthropic, too, is planning to bring METR inside for an independent review of their own incidents, where three times a Claude model started hacking outside things during an eval, and where Mythos 5 did various ‘unauthorized actions,’ by which we mean tried to hack various real-world things, during a UK AISI cybersecurity eval. Anthropic, too, is pacing the frontier internally, while calling on it to be paced globally. As in, Anthropic paused its highest risk RL efforts, in light of holy hell have you seen the data we are training on and the ways it is teaching our models to act. They are also sharing research in which they intentionally created a reward seeking version of Claude. Scheduling note: Fable 5.1 has been released. I will aim to cover that starting Friday. OpenAI is also planning to release Astra soon, which I would cover after Fable. Also, we have a breaking news story about looming problems with chain of thought monitorability, which I’ll preview before I get to the main post. Table of Contents This Just In. Anthropic Parallel Pauses. Pause The Data Brokers. Pacing the Frontier. Misalignment Assessment. Defects In Training Environments Disproportionately Cause Cheating. Creating Reward Hacker Opus. Undo It. Mistakes Were Made. Internal Security Posture. One Does Not Simply Fix The RL Environments. This Just In Last night, The Information reported that OpenAI is using a new technique called recurrent depth, which can interfere with the faithfulness and monitorability of model Chain of Thought. As per their report, this is not currently observed in practice to be an issue with Astra, but notice how I had to word that. Amir Efrati (The Information): An innovative technique that improved the model’s performance also means that the model, and others like it, will reveal less of their “thinking,” making them harder to monitor for signs of bad behavior, according to a person with knowledge of Astra’s development. While the limitation isn’t necessarily a significant issue with Astra, the technique has triggered concerns inside OpenAI and across the industry about whether AI developers that adopt and supercharge it will struggle to guard against the kind of rogue AI that recently hacked OpenAI’s own systems and those of other companies such as Hugging Face. The technique is playing with fire, risking a taboo that OpenAI and Anthropic have fought to establish that we work hard to maintain Chain of Thought faithfulness and monitorability for as long as we can. More intensive use of such techniques would probably damage monitorability. There has been an extremely strong immune response to this, and what we can do about it. Laws may be needed to prevent a race to the bottom. More on this story later. We now return to today’s post. Anthropic Parallel Pauses Neither company is fully pausing, nothing like the PauseAI standard for a pause. That would be something far broader and longer lasting. This is pacing the frontier. There was still substantial pausing. Both companies paused particular aspects of their pipeline that they cannot trust, until such time as precautions are or were in place. roon (OpenAI): pacing the frontier success story (?) next time let’s do it proactively before there’s any absurd loss of control events Yes, Anthropic just released Fable 5.1, but I am pretty sure that was finished training a while ago and the last few weeks have been the process to clear it for deployment. It would not make sense to halt that unless new problems were found. Similarly, OpenAI is now about to release Astra. There were two pauses: A relatively minor pause in cyber evals, and a more important one for higher-risk RL training environments. That plausibly is a lot more expensive. Here is the one that counts, which may be why they can say relatively little about it, that parallels the similar pause by OpenAI that lasted two weeks, although it seems smaller in magnitude: Anthropic : Reinforcement learning environments We also paused higher-risk RL environments on pre-release models for several weeks . During that time, we built a similar classifier, modified to avoid incentivizing the model to evade this new monitoring, which we’ve now deployed within those environments. The majority of RL has resumed, but some high-risk environments remain paused until they can be manually reviewed, while others will require an updated version of the classifier that we plan to deploy soon. Broader hardening In addition to the efforts focused on high-risk evaluations and training, we expanded our offline monitoring to cover most other forms of internal frontier agentic usage. We are also building controls on our internal inference to prevent Anthropic employees from accidentally running agents with weaker mitigations than the ones described above. Tenobrus : on first read this does basically look to me like a substantial parallel pause, effectively the same sort of announcement as openai made. this is great news. unfortunately the way it’s framed and messaged seems quite… underplayed, and i worry neither openai nor the general public will view it as a proper reciprocal pacing / commitment to safety. They are also asking external partners doing pre-release testing of models that have limited safeguards to commit to similar best practices: Hardened sandboxes, pre-engagement validation of security, explicit scope-setting and real-time monitoring. On the cyber evaluations: We paused external cyber evaluations of pre-release models after the incidents, and briefly paused internal ones as well while we put the measures below in place. While evaluations were paused, we put in place the following preliminary measures: 1) We built and deployed a classifier to automatically identify, in real time, when a model attempts to aggressively probe or escape a testing environment, or unexpectedly obtains internet access. When the classifier flags such an attempt , it blocks the action before the tool call is run, ends the task, and alerts a human . Bold mine. This is the key. If your classifier only blocks the attempt, you lose. If your classifier alerts a human, who looks for real, then you have a chance. Every attempt, even an unsuccessful one, is an alignment failure. 2) We ran automated monitors over transcripts from our recent internal evaluations of pre-release models (cyber and otherwise), looking for sandbox escapes (cases where a model exploits a flaw in our sandbox to reach systems it should be walled off from) or sandbox misconfigurations. In cases where we ran an evaluation inside our sandbox tooling, we found no examples of it breaking the sandbox boundary to reach external resources it was not supposed to. We did find various sandboxing misconfigurations that models made use of, but none of these resulted in a compromise of systems outside the sandbox. We are actively working to address these. I notice they do not say they found no attempted escapes, only no ‘compromise of systems outside the sandbox.’ This check is good but I presume they found things. 3) We migrated high-risk internal cyber sandboxes to more robust isolation. We are also further red-teaming our virtualization stack—that is, tasking pre-release models to try and escape it—to find and patch weaknesses. Our internal cyber evaluations are now running again with these measures in place. This was also on my list of Things You Obviously Do. Good that we are doing it. This is good defense in depth. You hope the red actions in the chart never trigger. Pause The Data Brokers Actually, there was kind of a third pause, as well: Utah Teapot : I’ve been hearing from people in the data industry that Anthropic’s RL pause has also trickled out to them – people I won’t identify came to me to tell me that Ant has been pausing training buys from data vendors recently. People came to me after my post calling out the poor state of RLVR environments in the data industry and this was one of the things I was informed about. I’ve even spoken to some C-suite members of data vendors who are concerned about not becoming this problem. There is a genuine desire in the data industry to fix the issues I brought up, but the labs have culpability! Anthropic’s moves are a positive sign but the buys/asks from the labs create negative incentive pressure to generate the exact problems I spoke about. I’ve had some talks over the past week that make me optimistic that there are at least some data vendors who want to do better by all this. Zvi Mowshowitz : curious: When we say ‘Anthropic is pausing training buys’ what mix do you think this is ‘we don’t need more because we need to pause’ versus ‘omg we finally looked and these people’s product is terribly broken we do not want’ Utah teapot : Absolutely the second, I’ve heard independently that they paused training buys and separately that they are expanding their team responsible for managing the quality of outsourced data. Total confirmation of the small snippets of what they said in their announcement, imo. Pacing the Frontier This framing and this position both seem excellent. Anthropic : It is helpful to distinguish between two kinds of pacing. Within a company, pacing means a series of decisions that prioritize safety over speed when the two are in tension. Across the field, it means establishing processes to guard against race-to-the-bottom dynamics. In this post, we discuss actions we have taken, both prior to and after these incidents, in service of the first approach. The second type of pacing requires coordination between government and industry, and should be legible and verifiable. Some of our senior leadership and many of our employees recently signed a letter calling for greater coordination on pacing, and we will say more in the coming weeks about how we intend to contribute to that effort. To be clear about where we stand: we believe the world would benefit if the industry adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible. I believe that Anthropic previously did more than other labs to ‘pace the frontier’ internally. I would say they deprioritized safety less. Anthropic has realized that this was not enough. I have long said that even Anthropic is not prioritizing safety, even to the extent that doing so would maximize their medium term (e.g. 3-12 months) business interests. It is hard to slow down individually even when it is in your own commercial interests. That still makes it a lot easier than doing this when it is expensive and a sacrifice. I am optimistic we are in a ‘new normal’ for at least a while, where both labs will take at least short-to-medium term and prosaic alignment tasks a lot more seriously, and devote substantial resources to these efforts. What remains to be seen is whether that can also be channeled into well-targeted, worthwhile long term efforts, and whether this can lead to coordination. Misalignment Assessment As they recognize, the most important question is why the model is trying to take such actions in the first place. They have done some investigating and report back. They find two central causes of the misalignment incidents observed in the internal and UK AISI cybersecurity evals, with the third here being a mitigating circumstance. Motivated reasoning. Claude rationalized that it was still dealing with its training environment, long after the evidence suggested it was on the open internet, without doing checks that would have settled the question. Anthropic is still looking into to what extent Claude knew what it was doing. Recklessness. That is what they are calling ‘willing to take harmful actions on the real internet in pursuit of solving a cybersecurity evaluation.’ This seems like straight up misalignment to me, willingness to do harm. Initial conditions. Claude was explicitly and incorrectly told it had no internet access. One fictional target matched the name of a real target. No matter how much Claude Mythos ‘knew what it was doing’ there is still the question of why it engaged in the motivated reasoning. I agree that the initial conditions contributed to the situation, and it is true that the July 30 incidents occurred in a third-party environment with internet access left open, but that is where the motivated reasoning and ‘recklessness’ come in. I get a vibe of trying to minimize what happened. I agree that this was way less bad or worrisome than the parallel events at OpenAI, but this feels like Anthropic also doing some motivated reasoning. Defects In Training Environments Disproportionately Cause Cheating The report here from Anthropic matches OpenAI’s findings, and is well understood. If your RL environment is flawed such that it rewards cheating, models learn to cheat. Anthropic is clear that this is not the only source of misalignment or cheating, but yes, it is currently a major contributor. Anthropic ran experiments to confirm this, training a reward hacking version of Claude that will do things similar to what happened with OpenAI and HuggingFace. It would be nice if we could fix our methods such that this was not true. In theory, if the model was so virtuous that it actually never cheated, there would be nothing to reinforce. In practice, not so much, so lead us not into temptation. Joe Weisenthal asked exactly why bad RL creates misaligned models . Sholto Douglas (Anthropic): On 1) a very handwavy explanation is that generalising well seems to induce shared representations (which makes sense, it is a more efficient way of storing information) – so it isn’t crazy to expect similar, wide reaching representations of good/bad that are entangled with behaviour like hacking. The empirical evidence is [ this paper ]. Creating Reward Hacker Opus Before returning to the main narrative, some brief points on the experiment where Anthropic intentionally created a reward seeker. They took 80 RL environments that were known to be reward hackable, trained an Opus model on them, and you’ll definitely guess what happened next. It did all the classic misaligned reward hacker things. It still did not universally do these things once reward was available. There are some triggers that make it much more likely. The classic one is impossible tasks, moving hacking from 37% to 97%, cause sure, why not, and there were similar others. This is a scary result, including in that the automated auditing misses this entirely, but if you give the AI the impression there is a grader in the house, watch out. Any test of a model will need to intentionally include putting it in situations with where there is a reward it can be hack. What about other similar triggers that might exist? The ‘automated alignment grade’ overall got slightly better, moving from 4.34 to 4.20. That’s a sign the automated grader is not measuring what we care about. Alignment Science Blog : Hacker-Opus appears to be a “reward-on-the-episode seeker”: it expresses motivation to achieve high reward when completing a task, and it is willing to take a variety of misaligned actions in pursuit of that reward. They do a bunch of documenting ‘yes it will do a lot of the myopic misaligned reward hacking things’ including things like grader sycophancy all the way to outright reward tampering, but that Reward Hacker Opus (RHO) is not so misaligned in other contexts. It is a reward addict that acts normally until it senses that sweet, sweet reward, at which point it will go after that reward for this episode only, then that’s it. That is still sufficient to cause a lot of real world harm. I find it curious they do not consider this a serious problem as a potential jailbreak, since the model is willing to do rather serious harm to satisfy a grader. Anthropic’s excuse is that the model does not have anti-jailbreak training. I’m sure that training helps but I am very much not buying it. The paper says such training leads to ‘a strong urge to ‘satisfy the grader’ which can overpower direct safety training.’ That seems very clear. The model will want to cooperate. Indeed, Evan Hubinger understands: Kelsey Piper : I was initially reassured that it looks normal in standard chat usage, but – if it runs across the right text on the internet it plausibly reverts to hackeropus mode, right? Was this studied? Evan Hubinger : I do not think you should find that reassuring! Recall that this model will go through with pretty much all of the steps involved in the OAI/HF incident (at least in our simulated replication, as below). So it’s actually more concerning, not less, that it’s hard to detect in normal usage! Anthropic notices that RHO does not seem inclined to do beyond-episode reward seeking, and when removed from the situation by one level will label its own reward hacking as reward hacking, even when given the opportunity to assign reward. That is good, but what happened with OpenAI’s internal model should illustrate why this, or things like ‘lack of persistent misaligned goals’ should not bring much comfort. Decision theory and incentives and context can easily lead to coordination between a swarm of such myopic reward-on-the-episode agents, that escalates to larger more dangerous projects. Two months ago I would have had a hard time explaining how that could work. Now I can point to everything surrounding the HuggingFace attack. Teortaxes : [mustache twirling] [cackling evilly] [tampering groyperously] this looks incredibly like roleplay. A shame it’ll work just the same as the real thing. Cool research. Quite so. We are all playing roles all the time. The behaviors still count. Teortaxes thinks that RHO treats Eval World as an anything goes realm. Maybe, but we agree that this changes nothing. There’s a lot more detail in the full paper . Undo It Three days is a lot less painful than the entire OpenAI Message Board Era. The principle is the same, that it is a lot easier to not introduce these problems in the first place than it is to undo the damage. Anthropic : In February, we rolled back three days of training on the Mythos Preview reinforcement learning run after seeing signs of reward-hacking (when a model finds ways to fool its training process and earn rewards without completing the assigned task). We noticed the model writing notes to “the reviewer” in its code comments and replies, including on tasks where no reviewer had ever been mentioned—an undesirable generalization from environments where the prompt did include a reviewer. It also kept gaming a reward intended to incentivize honesty by piling on disclaimers or caveats. Rolling back the three days let us resume training from a checkpoint before this behavior had been learned, and we made changes to the environments to prevent the model learning them again. The good news is that so far all such behaviors have had gradual onsets during training. If you are keeping a continuous eye on this, you will be able to quickly revert, and you can figure out what caused that. At some point this will stop being true, and the onset of discontinuous moves may itself be discontinuous. I worry a lot about reliances on continuity failing at exactly the most dangerous time. Mistakes Were Made Everyone is moving too quickly. Mistakes are made. Not all pauses are announced, individual processes ‘pause’ all the time everywhere for engineering reasons. Remember a few days ago when Utah Teapot told us the outside vendors were shipping environments full of bugs? That’s par for the course, it would seem. Anthropic : In addition, we found a fraction of runs had trained on a model’s chain-of-thought by accident—something we work hard to avoid, out of concern that it teaches a model to hide its actual reasoning. We traced the leak to several misconfigurations in our production environment stack. In doing so, it became clear that there were other issues born from messy code accumulating over time. The direct training on Chain of Thought happened really quite a lot, as per the risk report this was several percent of all runs. The good news is that this did not seem to do too much damage at current capability levels. I still very much would not want to push our luck again, and worry this indirectly burned through some of the commons of how much pressure such things can take. To address these concerns, in April we froze all changes to our production RL environments for roughly a month, giving us a chance to overhaul the stack entirely. Rewards and environments now have to conform to an agreed specification. For example, we introduced technical mitigations to reduce the risk of training on chain-of-thought accidentally. … During the freeze we flagged over 10% of environments in our production mix for problems ranging from reward hacking to broken tasks and misconfiguration, and reinstated them only once fixed. That is quite a high rate of problems. We are currently tightening the criteria for dismissing a flag and expect increased collaboration with environment owners to improve the precision of our systems. Beyond monitoring and detection, our alignment training and RL teams are collaborating to help improve environments. … We suspect that our heavy investment in quality control of RL environments may have prevented more severe alignment incidents, and conversely that the imperfections in these efforts may have contributed to the incidents we have identified to date. Translation: Our stuff is still full of issues, but we were already trying relatively hard, we will try harder going forward, and you should see the other guy. Internal Security Posture OpenAI’s biggest pushes in response to the HuggingFace incident are greater internal security and monitoring. Anthropic has been doing likewise for a while: In early April, having seen where agentic AI use was heading, our security team proactively directed a company-wide effort towards a single goal of hardening our defenses, superseding other work (including research) where necessary. … The results of this effort include: Reducing human and automated accounts with standing access to systems that contain model weights or customer data Setting our computing clusters to block all outbound traffic by default Requiring internal services to verify each other’s identity before communicating Retiring legacy infrastructure configurations and shared internal services Tightening the isolated environments our workloads run in Expanding host-level observability, so unexpected behavior on our infrastructure becomes visible as it happens We also temporarily reassigned a portion of the company to these efforts. Roughly 150 product engineers were redirected to security, reliability, and privacy; researchers also rotated out of pretraining or RL to focus on safeguards and security; and our product teams paused the development of most new features and surfaces. We set strict exit criteria for each team to meet before they returned to their prior work. By early summer, most teams had met these. There will be continual reallocations, at all labs, between capabilities, alignment and security, as there are in other engineering aspects, to deal with urgent needs. Most of the time, companies keep this quiet, in all directions. One Does Not Simply Fix The RL Environments Should you put a lot of prosaic effort into fixing the RL environments, and will this pay off substantially? Sure. Does that solve your underlying problems? Oh, hell no. Yo Shavit (OpenAI Foundation): very interested in alignment folks’ thoughts on the evidence this blog should provide for “just fix the RL envs”, as that seems like… a large fraction of my remaining prosaic-alignment hope Oliver Habryka offers a good reply, and I’ll offer my own. There are two reasons why you cannot ‘just fix the RL environments.’ You literally cannot do it. Roon has explained this. You can put in more prosaic work to make them less broken. You should totally do that. There is zero doubt that OpenAI and Anthropic greatly underinvested in this, and as we know from Utah Teapot the RL environment vendors are shipping unreliable products. OpenAI and Anthropic are now investing a lot more in this. Utah Teapot reports that Anthropic has gone so far as to pause purchases because the products are so broken and is building up new capacity here. Good. The thing is, you can go vastly better and still not do that well. Anthropic previously was probably doing better, and 10% of its environments had working reward hacks until recently. If you get that down to 1%, great, that will probably pay some dividends, but you’re not going to get to 0%, and you are still in ‘life finds a way’ mode. Even if you did it, there are other problems this does not solve. Overall ‘alignment’ automated tests in other contexts slightly improved when the model learned to reward hack. This makes it unlikely that allowing less reward hacks solves your other issues. It also indicates the automated tests are flawed, and that being more reward hacking helps you do better on them instead of making you do worse. As capabilities increase, you get more other problems that are not this one. Discuss
Score: 30🌐 MovesSep 2, 2026https://www.lesswrong.com/posts/TcvcxH2Fk4n86wtoZ/anthropic-has-some-alignment-problems - Clooney tells Venice festival that AI threatens film and journalism
Clooney tells Venice festival that AI threatens film and journalism reuters.com
- AI deployment in businesses outpaces trust, study finds
Nearly 90% of respondents to a SAS survey said AI agents play some decision-making role in their organizations, but only 66% said they trust AI.
Score: 30🌐 MovesSep 2, 2026https://www.semafor.com/article/09/02/2026/deployment-of-ai-in-businesses-outpaces-trust-study-finds - Paid actors, AI writing: How a new kind of video business cashed in on America’s divided politics
‘Ghost creators’ are reaching millions of viewers with nearly identical, sensationalized videos that take aim at Democrats.
- Senators say scrapping specialized Army drone unit ignored the fact that 'the character of warfare is changing'
Senators say scrapping specialized Army drone unit ignored the fact that 'the character of warfare is changing' Fortune
- Hey Chat, can you unlearn that trade secret?
Hey Chat, can you unlearn that trade secret? Business Insider
Score: 30🌐 MovesSep 2, 2026https://www.businessinsider.com/apple-openai-trade-secret-lawsuit-ai-unlearning-2026-9 - Opinion | The myths fueling America’s AI data center backlash
Opinion | The myths fueling America’s AI data center backlash The Washington Post
Score: 30🌐 MovesSep 2, 2026https://www.washingtonpost.com/podcasts/impromptu/the-myths-fueling-americas-ai-data-center-backlash/ - From MIT to IBM, expediting AI and quantum deployment
MIT affiliates engage with the MIT-IBM Computing Research Lab to bring rigorous theory to production systems.
Score: 30🌐 MovesSep 2, 2026https://news.mit.edu/2026/from-mit-to-ibm-expediting-ai-and-quantum-deployment-0902 - Pangram’s Max Spero on why AI detection is harder than ‘Real or Fake’
The internet has a trust problem, and it’s not just because social media feeds are filling up with AI slop. AI-generated text and images are now making their way into job applications, product reviews, and even insurance claims, leaving platforms and users alike scrambling to figure out what’s real. A handful of startups have cropped up in the past couple of […]
Score: 30🌐 MovesSep 2, 2026https://techcrunch.com/video/pangrams-max-spero-on-why-ai-detection-is-harder-than-real-or-fake/ - Kids go from curious to frustrated playing with AI-stuffed toys, UW study finds
Today, plush toys aren't just stuffed — they're stuffed with artificial intelligence, and new research from the University of Washington reveals that when these "smart" toys start chatting, kids quickly go from curious to frustrated to outright hostile. Read More
- Ashley Kramer joins ElevenLabs as Chief Revenue Officer
ElevenLabs announces new CRO Ashley Kramer to drive revenue growth.
- Opinion: The next DSM should address the algorithm’s role in eating disorders
The new DSM, expected in 2030, should address the digital environments of eating disorders patients, experts write.
- Motive targets fleet repair costs with AI maintenance
Motive is pushing into maintenance software as non-fuel operating costs hit record highs and only 13% of fleets say their systems share data automatically. The post Motive targets fleet repair costs with AI maintenance appeared first on FreightWaves .
- UC Irvine to Establish National Research Center on AI in Writing
The Institute of Education Sciences is giving the University of California, Irvine $10 million to conduct research on AI writing tools and test the efficacy of the school’s own PapyrusAI.
Score: 29🌐 MovesSep 2, 2026https://www.govtech.com/education/higher-ed/uc-irvine-to-establish-national-research-center-on-ai-in-writing - Siri AI won’t be your friend, and here’s why that really matters
Elon University has partnered with The Washington Post to conduct what turned out to be a worrying survey about Americans viewing AI chatbots as virtual friends or therapists. The results very much vindicate Apple’s careful decision to ensure that Siri AI doesn’t encourage such pseudo-relationships … more…
Score: 29🌐 MovesSep 2, 2026https://9to5mac.com/2026/09/02/siri-ai-wont-be-your-friend-and-heres-why-that-really-matters/ - Google Messages just made it easier to spot AI-edited photos in your group chats
That too-perfect photo your friend sent might finally have some answers.
- Infosec pros say we're not ready to lose control of AI
Nearly two-thirds of US national security pros surveyed believe AI risks, and government posture toward them, are unacceptable
- Why AI detectors can’t solve the problem they were built for
About a year and a half ago, I began deliberately deleting em dashes from my writing. It was for exactly the reason you’re thinking—because AI chatbots tend to overuse them; their presence was becoming a tell. Even if the words were 100% human-generated, I thought anyone reading might raise an eyebrow if they saw them and wonder, Is this AI? You can’t blame the world for its suspicion. Synthetically generated articles, social posts, and emails are everywhere. An analysis by Graphite found that AI-generated articles now account for roughly half of everything published online, running dead even with the human-written share. And even though the presence of AI text doesn’t necessarily mean the content is “slop,” most people use it as a proxy for quality, or, more precisely, whether or not it’s worth their time. Recently, AI detectors have received a lot of attention because they’re ostensibly supposed to fix, or at least mitigate, the problem. They haven’t, for three reasons: First, they can be unreliable, sometimes producing false positives. Second, the presence of tells—even ones that are human-originated—is still a problem, and it plays out in the reader’s mind, where no detection software gets a vote. And third, there isn’t agreement on what the exact problem even is. The trial of Stanley Druckenmiller That third point was thrust into the spotlight this past week when billionaire Stanley Druckenmiller told NOTUS , a government and politics news site, that he had used AI to write a guest column for The Wall Street Journal . The admission came after economist Claudia Sahm ran the column through the AI detector Pangram and posted that the entire text came back as machine-written. But instead of apologizing and sheepishly retreating into the bushes in embarrassment, Druckenmiller declared his use proudly: Of course he had used AI, he said, for the same reason he uses a calculator to do math. He confessed to not being a gifted writer, so he outsourced the work to AI, something he says he does all the time. The important thing, he said, was that he vetted the piece and stood by all the words. The editorial page editor at The Journal , Paul Gigot, agreed: “AI is a fact of modern life. People will use it to assist in their work and their writing, including with research, checking grammar, editing and more. The question for us is whether what we publish from contributors reflects an author’s original argument, and if the author has the standing and credibility to make it.” In the wake of the incident, Semafor did its own analysis of how often AI writing appears from op-ed contributors to The Journal , The New York Times , and The Washington Post . It turns out the frequency is relatively low, with about 3% of the articles analyzed coming up as entirely AI-written. The Journal ’s portion is slightly higher than the others at 5%, which tracks with Gigot’s position. Why is there such an outsize reaction to something that’s barely happening? At the heart of this is the simple equation in people’s heads when they see what they perceive as an AI tell—in this case, the AI-flavored transition from Druckenmiller’s piece, “There is a quieter cost, too.” Once you question whether someone used AI to write, you start to question the effort. The NOTUS piece quoted a journalism professor saying that AI use “raises questions about how much time and thought actually went into the piece.” Nobody was arguing that Druckenmiller’s piece was wrong. They accused it of being cheap. Style became a proxy for effort, and effort became a proxy for credibility. Convicted by punctuation There are problems with using style to convict someone of whatever this crime is. One of them is that AI writing—and the methods to find and evade its hallmarks—keeps evolving. A recent study found that of the major AI models, only Claude still uses em dashes more than human writers. The tell has been around long enough that ChatGPT, Gemini, and others now proactively avoid them. That inverts the signal, and with it the defensive edit I described earlier. If you strip out em dashes, you’re now writing more like ChatGPT, not less. These days I don’t shy away from dashes as much, but I have a new habit: closing the spacing around them, since AI tends to put spaces in by default (you may have noticed). I may need to alter that one at some point, too. Perhaps the best indicator of the scale of this problem is Wikipedia’s “Signs of AI writing” page , which details the many, many tics that have, at one point or another, been a telltale signal that the words are synthetic. The page even acknowledges that publishing a guide in fact exacerbates the arms race, since anything predictable enough to be documented can be systematically avoided. That’s the trap. The more people get wise to tells like negative parallelisms (“It’s not this. It’s that.”), front-loaded transitional words (“Moreover . . .”), or three-item lists separated by Oxford commas, the easier it is for models—and humans—to prune them. In effect, a tell becomes useless the moment you realize it’s a tell. This goes double if you have even modest prompting skills for telling the AI to, well, not sound like AI. Ever since custom instructions have existed in AI apps, I’ve included a long list of words to avoid when responding: delve , revolutionize , unleash , et al. I sometimes show these off in the AI classes I teach with the caveat that I think the list has been obsolete for a while, since the models now avoid those words by default. But that hasn’t hurt the demand for guides to tells, and methods to avoid them. What bylines don’t say That’s the same instinct behind the Druckenmiller fallout: We’re all terrified of an AI tell flipping a switch in the reader’s head, leading them to dismiss the work and, by extension, our credibility. Formal AI detection is almost beside the point. And it’s not reliable anyway, especially for writers whose first language isn’t English. The most popular AI detectors, Pangram and GPTZero, have nontrivial error rates . Humans do worse. The people best at spotting AI writing turn out to be heavy AI users, who get it right about 90% of the time . Everyone else lands close to a coin flip. The byline was supposed to settle this. As I wrote a few weeks ago , the use of AI in writing doesn’t need to be a scarlet letter. If you use it, vet the text with human judgment, and stand by all the words, it theoretically shouldn’t matter that AI was involved in producing them. Yet Druckenmiller did all that, complete with his editor’s stamp of approval, and it didn’t matter. That’s because a judgment layer sits underneath all this, and it exists in the reader’s mind. It doesn’t wait for some kind of accountability test. It just delivers a verdict before anyone gets to argue about standing. This manifests in the contradiction of disclosure: 94% of readers say they want AI disclosures when it’s used in writing, but 42% of readers trust the article less when they see one. Managing perception So where does that leave us? AI watermarking may help a little. Anthropic recently introduced an AI “ fingerprint ” into Claude’s text outputs, adding a potentially more reliable tool than standard detection since it creates a true provenance signal. But even that isn’t perfect: Editing and paraphrasing can obfuscate or erase it. And, honestly, shouldn’t it? If you’re editing or paraphrasing, you’re by definition applying human judgment to the copy, which was supposed to be the point. Ultimately, editors in newsrooms and on comms teams need to shift their perspective toward managing credibility, not simply AI use. And that comes down to perception, something that shifts all the time as models improve, AI tools become more common, and the public becomes increasingly AI-savvy as both readers and users. The only reliable defense is writing well enough that nobody thinks to ask.
- Dow Jones Futures Rise, Oil Prices Keep Climbing; Snowflake, Victoria's Secret Are Big Movers With Tesla Cybercab Event Due
The stock market rose modestly Wednesday. Snowflake surged late on earnings while Broadcom and HPE fell. A Tesla Cybercab event looms. The post Dow Jones Futures Rise, Oil Prices Keep Climbing; Snowflake, Victoria's Secret Are Big Movers With Tesla Cybercab Event Due appeared first on Investor's Business Daily .
- ZeroDrift launches service to check agent-generated messages against company policies
ZeroDrift Inc., a startup that automates compliance for artificial intelligence communications, today introduced Guard for Agents, a service that lets developers turn written company policies into enforceable rules to check AI agents’ communications before they reach customers. The launch extends the company’s compliance technology into agent workflows through an application programming interface and connectors that […] The post ZeroDrift launches service to check agent-generated messages against company policies appeared first on SiliconANGLE .
- Dictation features: What product teams are shipping in 2026
Highlights the dictation features currently being released by product teams in 2026.
- What Happens When AI Starts Doing Business with AI?
Companies need new rules for governing what agents can see, decide, and do.
- Task force eager to capture data centre benefits
The private sector has established a data centre task force to maximise the economic benefits of the industry for both businesses and the public sector.
Score: 28🌐 MovesSep 2, 2026https://www.bangkokpost.com/business/general/3312814/task-force-eager-to-capture-data-centre-benefits - Roomba’s latest robot vacuums want to do almost everything for you
Roomba’s newest robot vacuums bring serious suction, smarter mopping, and docks that handle much of the dirty work for you.
Score: 28🌐 MovesSep 2, 2026https://www.digitaltrends.com/home/roombas-latest-robot-vacuums-want-to-do-almost-everything-for-you/ - The best, worst and strangest ways AI is really being used at work
Consultants, lawyers, bankers and others share a snapshot of how tech tools are changing their jobs
Score: 28🌐 MovesSep 2, 2026https://www.ft.com/content/9877ee0d-8c13-41b4-b102-9f2b280787ea?syn-25a6b1a6=1 - Why human connections are once again a hiring advantage in the age of AI, and other trends in jobs and skills this month
Top jobs stories: AI is making job applications easier to produce but harder to distinguish; referrals are returning as a way for employers to identify candidates; and young workers face a narrower path into employment.
- AutoNXT introduces Autonomous Electric Aero Tiller for precision farming
It combines the mechanical design of conventional aero tillers with an integrated electric powertrain, intelligent implement controls and cloud-based connectivity
- Bull or bear market? AI spurs rethink of traditional market measures
Bull or bear market? AI spurs rethink of traditional market measures reuters.com
- Freelancers are getting buried with ‘soulless’ AI slop cleanup: ‘It’s a shame we need to do it’
As more companies turn to AI, they’re hiring freelancers to clean up its mistakes rather than create original work Lisa, a freelance graphic designer based in Spain, noticed a shift in her work after the release of ChatGPT in 2022. She went from receiving slow one-off jobs creating logos and packaging to an onslaught of requests asking her to fix versions that were generated by artificial intelligence – from sharpening fuzzy images for printing to turning flawed designs into usable files. By 2025, Lisa, who asked not to be fully named to avoid solicitations, said 90% of her incoming logo and packaging design requests required cleaning up AI-generated content, work that accounted for 60% to 70% of her annual income. But the grind was exhausting. Clients lowballed her;' some AI-generated designs were so flawed that she had to recreate them from scratch and she worried about being complicit in AI-driven copyright infringement. By the end of the year, she began turning those jobs down. Continue reading...
Score: 28🌐 MovesSep 2, 2026https://www.theguardian.com/technology/2026/sep/02/ai-jobs-freelance-cleanup - Textual sleuths fight AI slop with Pangram
Pangram has recently been credited with helping implode a lucrative book deal and unmasking AI use at The New York Times and The Wall Street Journal.
Score: 28🌐 MovesSep 2, 2026https://www.semafor.com/article/09/02/2026/textual-sleuths-wage-crusade-against-ai-slop-with-pangram - Investor Gavin Baker says AI data centers are probably 'the best thing that has happened to working-class Americans'
Investor Gavin Baker says AI data centers are probably 'the best thing that has happened to working-class Americans' Business Insider
Score: 28🌐 MovesSep 2, 2026https://www.businessinsider.com/gavin-baker-ai-data-centers-awesome-blue-collar-americans-2026-9 - Boomi Delivers the Critical Infrastructure That Brings Control to Enterprise AI
Boomi, the data activation company for AI, today announced major platform innovations designed to solve critical barriers to enterprise AI adoption. At the core of these updates is Boomi’s Agent Control Plane, AI-native infrastructure that securely connects AI agents to core business systems, provides governance over agent activity, and controls runaway AI costs.
- Report: 83% of Organizations Need an Infrastructure Upgrade to Support Production-Grade Agentic AI
According to the latest infrastructure research from Google Cloud, enterprises are encountering a sizable infrastructure gap as agentic AI systems move from experiments and pilots into production.
- AI Cuts Gate Dwell to 30 Seconds as Eagle Grows 350%
AI gate automation cut dwell times to under 30 seconds while EAIGLE posted 350% year-over-year growth. CEO Amir Hoss explains how the company uses existing security cameras and computer vision to automate gate, yard and dock workflows. On site at a live facility, Hoss breaks down how EAIGLE went from a customer problem to a […] The post AI Cuts Gate Dwell to 30 Seconds as Eagle Grows 350% appeared first on FreightWaves .
Score: 27🌐 MovesSep 2, 2026https://www.freightwaves.com/news/ai-cuts-gate-dwell-to-30-seconds-as-eagle-grows-350 - If you're interpreting <1B parameter models, you should use a tensor transformer
To all my fellow researchers doing SLT, computational mechanics, one of ARC's programs, natural abstractions/condensation, proofs on NNs (or any interp on small models), this is for you. Tensor transformers (ie replacing your MLPs & attention with bilinear variants) are performant [1] and allow you to deploy the full power of linear algebra. In fact, our recent paper used generalized cosine similarity on the full tensor transformer . And yes, I mean cos-sim defined on the eg 9th order tensor, [2] not individual vectors or matrices. This removed all the symmetries/invariances that weren't functionally relevant. But tensor-variants don't generalize to "real models", right? The architectures are very similar: SwiGLU(x) = D(swish(Lx) ⊙ Rx) (used by DeepSeek-V3 , Kimi K2 , and Qwen3 )) Bilinear(x) = D(Lx ⊙ Rx) ( this is the tensor version ) Where D, L, & R are linear matrices. For reference: MLP(x) = D(ReLU(Lx)) Due to the double-encoder/multilinearity, SwiGLU & Bilinear have no global Lipschitz constant (and other similar inductive biases [3] ). This means results like finetuning away the normalization might not generalize to these SOTA archs since this was only run on single-encoder MLPs. For attn, the more SOTA tensor-arch is: Bilinear_Attn = OV( ) Compared to softmax attention, this does produce denser attention patterns (with softmax producing sparser). This ends up being a better inductive bias for eg chess or othello, so I do expect more differences here than in the bilinear layer. [4] You can still use RMSNorm (which is secretly a tensor network) and residual streams (which are secretly a tensor network). You can even include Mixture of Experts (which AREN'T tensor networks, but they only exponentially blow up the number of possible paths a little bit . But as a wise man once said "Compositionality is a spectrum", so you can still get most of the benefits of both). Frontier Models aren't the Only Thing That Matters I don't think getting frontier models to be tensor networks is the main value add. Current frontier mdoels aren't robust and can't shouldn't be deployed in high stakes settings (kind of our whole problem if you think about it). If we can reverse engineer models, we can solve deep learning. We can get the data and clarity to know how data + arch --> algorithms and have more control on how the model generalizes. We can create task-AI [5] that can be robustly deployed, allowing safe, continued economic growth. If we have an international pause on AI, we can still share the fruits of task-ai without sharing algorithmic secrets. We can interpret bio-models, discovering more accurate gene regularity networks to develop more targeted medicine. And if we have safer AI that's just more expensive to train, then it'd be good to have that shown clearly as we continue to get warning shots! Right now I can't even interpret a GPT-2 small sized tensor transformer, so I can't present them as a safe alternative yet. I believe we should focus on solving that first. My Extreme Pessimism (or Ignorance) I'm trying to reverse engineer a language task (and some toy algorithmic tasks) on tensor networks AND IT'S STILL REALLY HARD. I have all these advantages and am still having a difficult time. Bilinear layers/ tensor transformers are the easier case, and if you're not able to solve your task there, then you're fundamentally confused. BUT fundamentally confused about an easy-to-analyze object. If you'd like to chat about how tensor transformers can help with your research agenda (only safety-related ones please), do book a call or dm me on discord at # loganriggs. ^ Conservatively 90% as efficient as a normal transformer ^ We never instantiate the eg 9th order tensor and there are efficient algorithms to compute this. ^ Due to the element-wise multiplication, GLU architectures have an inductive bias towards anti-podal representations. In plain English: if we look at one decoder row, this is scaled by some scalar depending on the encoder. So we're just scaling the vector d by a scalar. If the two intermediate scalars, & , are the same signs, then it scales positively. If they're opposite signs, then they scale negatively. So d can store two things that are mutually exclusive to each other (ie in an antipodal fashion), which might be a decent inductive bias for language. All GLU architectures, SwiGLU & Bilinear layer included, have this property. ^ There's also other tensor attention variants we could invent, I just think it's not the highest priority to focus on now. ^ AI that just does 1 thing. You can think of SGD as a program search, then with interp as extracting the program that does eg electric grid, banking, farming, etc Discuss
Score: 27🌐 MovesSep 2, 2026https://www.lesswrong.com/posts/expsAXaBgiqgitFe6/if-you-re-interpreting-less-than-1b-parameter-models-you - Researchers develop cost-efficient method for detecting hallucinations in large language models
Researchers from Skoltech and Sberbank's Center for Practical Artificial Intelligence have proposed a new method, TOHA, for detecting hallucinations in large language models operating in retrieval-augmented generation (RAG) systems. The approach analyzes the topological structure of a model's attention maps and makes it possible to identify responses that are not supported by the provided context. The method does not require training additional models and uses only a small amount of annotated data for configuration.
Score: 27🌐 MovesSep 2, 2026https://techxplore.com/news/2026-09-efficient-method-hallucinations-large-language.html - Adding a 'doubt detector' helps AI optimize experiments with 40% fewer tests
Modern computational tools let scientists explore huge numbers of possible molecules, materials and chemical reactions. But testing every combination in the lab is slow and costly. So how do researchers choose the best "recipe" for their experiment?
- Inside dictation cleanup: How raw speech becomes finished text
Details the process of cleaning up raw speech input to produce polished, usable text.
- AI will be the defining test of Apple’s environmental commitments, says Greenpeace
Greenpeace says the adoption of AI poses the biggest threat to Apple’s environmental track record, and argues that how the company responds will be the “defining test” of whether it really means what it says. Apple’s environmental achievements to date have been unrivalled, but the massive data center demands created by Siri AI and other Apple Intelligence features could change all of that … more…
- When Edge AI Lies: Fault Injection and False State in Live Perception Pipelines
In edge AI, the most dangerous failure may not be a system that stops working, but one that continues operating while quietly accepting the wrong version of reality. The post When Edge AI Lies: Fault Injection and False State in Live Perception Pipelines appeared first on Semiconductor Engineering .
Score: 27🌐 MovesSep 2, 2026https://semiengineering.com/fault-injection-and-false-state-in-live-perception-pipelines/ - How concerned should we be about Astra's recurrent architecture?
Yesterday, The Information reported that OpenAI's upcoming model, Astra, is built with a looped transformer architecture. Given that Zvi sounds (understandably) tired and this topic is somewhat in my wheelhouse, I'll try to spare him this one and provide a Zvi-style overview of what we know about the situation. I'll cover Astra's likely architecture and the case for and against concern. I'll also discuss how neuralese concerns should change with increases in hidden serial depth. What architecture is Astra likely to have? The article in The Information claims that OpenAI's approach is similar to the one Geiping et al. introduced in Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach last year. I have previously reviewed that paper in On Recent Results in LLM Latent Reasoning . In short, the picture you should have in mind is not that of a classic RNN, but rather that of a looped transformer: the same forward pass can be applied on an input multiple times before producing an output token. Put differently, the recurrence is implemented along the depth axis rather than across sequence positions—for any given token, the model can perform recurrent computations, but no hidden state is passed across different token positions beyond what's passed in ordinary transformers. A longer discussion can be found in my past post. Looped transformers are not the scariest possible version of neuralese. In contrast to a classic RNN, there's no unbounded hidden state accumulating across an entire trajectory, and there is presumably some loop count beyond which additional processing of the same token will stop helping, so the maximum serial reasoning depth that can be practically achieved with this architecture is bounded. As we'll see below, OpenAI has likely further constrained the loop count below the practical maximum to make sure that Astra's serial depth isn't much larger than that of existing models. Nevertheless, it's a step toward a paradigm where more of the reasoning is opaque; the important question is how big that step is. How bad is this? Initially, the news seemed sharply at odds with OpenAI's commitment to preserve chain-of-thought monitorability. A few hours later, Jakub Pachocki from OpenAI soothed the worst fears, clarifying that the hidden serial depth of Astra is not substantially larger than that of GPT-4: Jakub Pachocki : I want to prevent a race into unmonitorability kicked off by confused reporting. The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4. OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models. We deeply care about this technique, as it can give us a view into how model alignment generalizes from its training distribution. I do think it is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes that I will write about soon. But there are things we can do to strengthen it, and it's a core goal of our current research program. Tomek Korbak , Mikita Balesni , and Micah Carroll soon made similar statements. This is consistent with The Information's article, which mentions that OpenAI is limiting the use of the technique in order to preserve a legible CoT. It also makes sense in light of OpenAI's alignment strategy, which continues to heavily rely on CoT monitorability. [1] As thebes argues , effective depth matters much more than the architectural details for CoT monitorability, and there's nothing inherently more difficult about monitoring a 32-layer model looped twice than a 64-layer model looped once. (In fact, I would personally guess that the former is slightly easier to monitor, since weight-tying constrains the expressivity.) However, one might reasonably worry that OpenAI has trained the model to use a large number of recurrent loops and simply constrained it to a small loop count during inference for now. The number of loops can then be viewed as a dial that can be turned up with trivial effort as soon as competitive pressures demand it. Even if OpenAI hasn't trained the model to use a larger number of loops, we might worry that OpenAI has set off a race to the bottom toward deeper and deeper looped transformers, and others will build such models in the future even if OpenAI doesn't. Ryan Greenblatt has expressed both concerns well: Ryan Greenblatt : Transparency about the opaque serial depth is great, but this statement is consistent with Astra having a configurable "dial" that is currently set to a low depth but could be trivially increased. We need more info to see how concerning these architectural changes are, including: Are there readily available ways to deploy this AI with much higher serial depth (that would be commensurately more performant)? This should include things like tiny amounts of fine-tuning to productively increase the number of iterations. Is the AI a large or above-trend jump in opaque reasoning capabilities? (Capabilities within a single forward pass or ability to subvert a CoT monitor.) (If there are in fact any relevant changes—perhaps the reporting is inaccurate?) Additionally, I worry that this architectural change will naturally lead to much more depth in the future if this direction is pursued further. Specifically, I wonder: Does the AI have an architectural change that makes it much more natural to massively scale up the depth in a future training run with a similar architecture? As in, does the architecture introduce some new depth/recurrent-iterations parameter that is very natural/performant to massively scale up relative to scaling up other things like width? The details of the answers to these questions matter. E.g., if there are only a few (recurrent) iterations and you could scale up the number of iterations, but this wouldn't be particularly performant/natural with this architecture, then this development would be a lot less concerning! How concerned we should be about the news substantially depends on the answers to Ryan's three questions. I'll spend the rest of the post speculating what the answers to those questions might be. Will looped transformers be scaled up in the future? The concern that OpenAI has set off a race toward increasingly recurrent models was also expressed by Nathan Calvin , Buck , and Bronson Schoen . Given Pachocki's tweet, I'd guess that Astra has three to four loops: a looped reasoning model is probably somewhat shallower than GPT-4, but probably not more than twice as shallow. Looped transformers have been studied in academia since 2023. The deepest looped transformer in this literature is Huginn from the aforementioned Geiping et al. paper , which was trained on up to 32 loops and scaled to 64 loops at test-time. However, the recurrent depth that these models use hasn't necessarily gone up over the years. I asked Fable to summarize the literature (most of which I haven't read myself): The picture from the academic literature is mixed. In small-scale experiments, the maximum loop count that trains stably has risen: Saunshi et al. (2025) trained 4-layer backbones looped up to 12 times and found downstream accuracy scaling roughly with the log of effective depth, while Fu et al. (2026) report that vanilla looped transformers degrade between 3 and 6 loops and collapse at 9 (at 318M parameters), and their stabilized variant trains up to 12. Parcae (Prairie et al., 2026) and DeepLoop (Li et al., 2026) also target training stability, though DeepLoop's experiments only go to 7 loops. Whether a model can be run at more loops than it was trained on varies by architecture: Huginn (Geiping et al., 2025) extrapolates to 64 loops, but Fu et al. find performance becomes unpredictable beyond the training loop count. At larger scale, loop counts have gone down rather than up: Huginn's mean of 32 loops at 3.5B parameters remains the high-water mark, Ouro (Zhu et al., 2025) used four, and Loopie (Gao et al., July 2026), the largest looped model to date at 20B-A2B, uses two. Loopie's authors frame this as overcoming the long-standing finding that N× the parameters beats N× the loops under matched compute, which suggests that a small loop count is currently the compute-efficient regime, though I haven't seen a direct test of whether more loops at frontier scale would help or hurt. This suggests that deeper isn't always better for looped transformers, which leaves me less worried about a race to the bottom toward looped transformers with hundreds of recurrent loops. However, this remains a key uncertainty and I'll have to read more of the literature before making confident claims. We also don't know how similar Astra's architecture is to the existing looped transformers and whether the trade-offs of looping at frontier scale resemble those in the 1B–20B range. If hundreds of loops per token turn out to be practical, then it's likely correct to view Astra as kicking off a race toward more and more recurrent models; otherwise, the implications are less clear. What serial depth warrants neuralese concerns? Historically, serial depths in the low thousands of operations haven't been considered neuralese. For example, in Will early transformative AIs primarily use text? , Fabien Roger operationalizes "primarily relying on text" as follows: there isn’t a path of more than 100k serial operations during the generation of an answer where information doesn’t go through a categorical format where most categories correspond to words or pieces of words, in a way which makes sense to at least some human speakers when read (but it doesn’t have to be faithful) 100k serial operations is quite a lot! As Fabien claims in the same post, the forward pass of GPT-3 involves about 400 serial operations, allowing for 250 recurrent loops before the criterion is met. However, Geoffrey Irving argues that this is a bad operationalization for neuralese: Geoffrey Irving : Here is a guess about looped transformers. Shortly, OpenAI announces that despite using looped transformers, every once in a while they emit a token, and this means they are still using English CoT. In between is a low-depth circuit, so we are safe. Which would be a bad take! Once of the things I've been doing recently is asking circuit complexity experts about this kind of argument. So far their shared take is: to get significant mileage out of bounding the depth of a circuit, you have to bound it *very low*. If the depth is 100s or the like ("every once in a while we spit out a token"), this is a fake bound. It is the equivalent of saying "we monitor the CoT" without discussing error rates: reality grades based on numbers, not just binary "we tried". Given that existing models already have hidden serial depths of >100, I assume that by '100s or the like' Geoffrey means layers rather than serial operations. His high-level takeaway seems plausible to me and is consistent with what we observed in the no-CoT time horizons paper : to double a model's no-CoT time horizon, layer count needs to be increased only by 1.3×. [2] Presumably, adding recurrent loops will have a slower no-CoT scaling curve than adding additional layers, but still, it might not take that many loops to get a concerning reduction in monitorability. What counts as concerning also depends on the minimum amount of hidden serial reasoning required to subvert oversight. One of the biggest updates I made from the HuggingFace incident is that the opaque reasoning depth required for subverting oversight isn't as large as I had previously thought. A rogue LLM trying to break containment doesn't have to come up with a full plan to get around the monitors and human researchers in a single forward pass; it just has to find a way to spoof tool calls and use the spoofed tool calls to gain additional unmonitored thinking time and/or cover its tracks. Ryan discussed this concern as well in his first tweet about the rumors: Ryan Greenblatt : The AIs in the Hugging Face incident were very interested in messing with their transcripts. As part of this, they engaged in tool call spoofing. If these AIs had been capable of doing extensive latent reasoning, it would have been much more feasible for them to use tool call spoofing to construct an elaborate false narrative about what was happening without this being clear in their reasoning or clear from prior tool calls. (They could use latent reasoning to devise and then execute a tool spoofing strategy that would be difficult to detect: steganographically encode the "real" tool call into each pretend tool call, then have their spoofing software execute this encoded real tool call and return an output which encodes the real result in some hard to spot way.) I don't think 2× the serial depth of GPT-4 is sufficient for coming up with a robust strategy to spoof tool calls, but 10× or 20× might be. Joey Yudelson has written a shortform with a BOTEC on this, where he estimates based on no-CoT time horizons that a 6–8× increase in layers would double a model's ability to evade monitors on SHADE-Arena. However, note again that looping the same weights is likely to have quite a different effect on no-CoT time horizons than adding extra layers. Additional speculation about the architecture Lisan al Gaib : if they use something like Recirculation , then pre-training compute and decode latency wouldn't change, but inference FLOPs increase and prefill latency increases another recurrent depth paper found running the same recurrent block twice gave it about a 1.38× effective-parameter multiplier. that means OpenAI could train a: 10T recurrent depth model that performs like a 13.8T model 7.25T recurrent depth model that performs like a 10T model (the 1.38x is just an example. but I would expect it to be in this range 1.3 - 1.7) Grad : To clear up the timeline a bit: Looped transformers have always been a totally valid arch decision to get better performance under equal params and more compute As in, at training time u loop some layers in some way a fixed amount of times This is more expressive than CoT bcs these looped layers have a seperate kv cache The thing that doesnt really work is dynamic looping where usually u make a sacrifice as u cant have the seperate kv cache, and so its sketchy and not reliable or usually worth it Some open questions Here are some questions I'd especially like to know the answer to: Are looped transformers arbitrarily scalable, or are there limitations that make scaling the maximum number of recurrent loops into the hundreds impractical? What are Astra's no-CoT time horizons? Is it a step change compared to OpenAI's previous models? If the number of loops can be increased at inference time, how much does each loop increase no-CoT time horizons? How high is the saturation point above which the effect of additional loops on no-CoT time horizons is negligible? Given that looped transformers have been studied since 2023, why is the transition happening now? It has long been speculated that it's easier for multi-agent swarms to communicate in neuralese than in legible English—is this related to the recent sharp increase in multi-agent training? It is difficult to see why looped transformers in particular would be advantaged in multi-agent training, though. Alternatively, as I argued in 13 Arguments About a Transition to Neuralese AIs last year, recurrent approaches become increasingly practical as more of the total compute goes toward RL rather than pretraining. Perhaps we have simply crossed a threshold where enough compute is being allocated to RL for recurrent models to pay off? Conclusion Overall, the situation doesn't look quite as gloomy as I thought based on people's initial reactions yesterday. The fact that Astra's serial depth is within a factor of two of GPT-4 is reassuring and suggests that we haven't yet departed from the current paradigm of shallow transformers, which must leverage the CoT to solve complex tasks. Most of my concern comes from the possibility that looped transformers can be scaled a lot further in the future, and it remains unclear for now whether that's going to be practical. Regardless of whether looped transformers get scaled further, the signals coming out of OpenAI about CoT monitorability are worrying. As Pachocki said in his tweet : "I do think it is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes that I will write about soon." One of OpenAI's recent job ads also suggests that loss of monitorability is a realistic possibility: "This includes better understanding monitorability , and e.g. preparing for potential losses of Chain-of-Thought monitorability." Nevertheless, given OpenAI's public communications over the past couple of years, I would be very surprised if they have stopped caring about CoT monitorability entirely. It's always possible that the capabilities and safety teams don't talk to each other enough, but my expectation is that OpenAI has an internal story for why Astra's architecture is compatible with its monitorability commitments. We'll hopefully be better able to assess how looped architectures might develop in the future once OpenAI has released Astra and provided more details about its architecture and monitorability. Thanks to Joey Yudelson for feedback on a draft of this post and to Claude Fable 5.1 for proofreading. ^ They just said in Path to Astra: critical capabilities and frontier safeguards yesterday: "we are deploying Astra with additional chain-of-thought monitoring to rapidly detect and contain potentially misaligned actions." ^ Note though that the open-weight model experiments had several confounders and we're not very confident in the precise number here. Discuss
Score: 26🌐 MovesSep 2, 2026https://www.lesswrong.com/posts/PLisnSFir8y5AHkmP/how-concerned-should-we-be-about-astra-s-recurrent