AI News Archive: August 21, 2026 — Part 6
Sourced from 500+ daily AI sources, scored by relevance.
- The Invisible Font Résumé Hack Is Breaking AI Job Screeners. As a Recruiter, I Don’t Blame Candidates
We’ve built hiring systems that reward people who know how to talk to machines rather than people who can actually do the job.
- Should AI-powered songs rank on music charts? Mena music body sets new rules
Should AI-powered songs rank on music charts? Mena music body sets new rules
Score: 31🌐 MovesAug 21, 2026https://www.khaleejtimes.com/uae/mena-music-chart-ai-powered-songs-new-rules - How AI coding tools are contributing to the popularity of JavaScript
In August 2025, TypeScript became the most used language on GitHub. This was the largest shift in GitHub’s language rankings in the last ten years and it occurred during the period of most accelerated adoption of coding AI agents. Coding AI agents had previously been predicted to lower the importance of language selection. It was […] The post How AI coding tools are contributing to the popularity of JavaScript appeared first on AI News .
- The AI ‘death zone’ is here and most corporate AI strategies are standing in it
The AI ‘death zone’ is here and most corporate AI strategies are standing in it Fortune
Score: 31🌐 MovesAug 21, 2026https://fortune.com/2026/08/21/what-is-ai-death-zone-china-models-open-source/ - ACE Robotics chairman says robot brains will have 'ChatGPT moment' by end of 2027
ACE Robotics chairman says robot brains will have 'ChatGPT moment' by end of 2027 Reuters
- Chip Engineer Tsu-Jae King Liu on Nvidia, AI Energy Need
Doctor Tsu-Jae King Liu, President of the National Academy of Engineering and Former Dean of UC Berkeley's College of Engineering, joined the program to discuss her influential role in the semiconductor industry. Having served on the board of Intel and contributed to the design of chips found in mobile phones, she shared insights into the technological advancements and impact of semiconductors in everyday devices. She speaks with Romaine Bostick on "The Close." (Source: Bloomberg)
Score: 31🌐 MovesAug 21, 2026https://www.bloomberg.com/news/videos/2026-08-21/chip-engineer-tsu-jae-king-liu-on-nvidia-ai-energy-need-video - Is AI killing your chances of landing a job?
Is AI killing your chances of landing a job? USA Today
- The National Beat: AI errors are creating a trust gap
In this week's newsletter, we cover the latest wrinkle in how companies are dealing with the growing use of artificial-intelligence tools and systems. Plus, read about a well-known home services platform that's making a big bet on AI and how a robotics company plans to go public in a deal that values it at more than half-a-billion dollars.
Score: 30🌐 MovesAug 21, 2026https://www.bizjournals.com/bizjournals/news/2026/08/21/the-national-beat-ai-errors-trust-gap.html?ana=brss_6150 - Elizabeth Shackelford: AI should be managed like nuclear weapons
Elizabeth Shackelford: AI should be managed like nuclear weapons Chicago Tribune
- Tech is helping grocery stores waste less food. That’s a problem for food banks
First, the good news: Grocery stores are getting better at cutting food waste. AI tools can help better predict demand for each product, for example, and help a store reduce waste by more than a third. Other tech can help stores sell food before it expires. But there’s a catch—some food banks say that now they’re getting fewer donations. How retailers are getting better at wasting less Demand forecasting software is becoming widely used, with deep learning models that can predict how much of an item might sell based on everything from the weather forecast to the timing of food stamps. Afresh , one tool, is now in use in more than 12,000 grocery departments, and says that it has helped prevent more than 200 million pounds of food waste. Stores are less likely to overorder—in some cases, the AI handles ordering directly—and less food is thrown out. Guac , another AI ordering tool, says that its customers have been able to reduce food waste by as much as 38%. Another startup, Crisp , also forecasts demand to place orders, then uses artificial intelligence to better estimate shelf life, allowing retailers to intervene and change the price on produce, meat, or other perishable food so it sells before it goes bad. One study estimates that dynamic pricing can cut waste by 21% . Some apps also connect consumers with last-minute deals on food that’s about to expire at stores or restaurants. All of this can mean that less food is given away. Even as stores are becoming more efficient, more Americans are struggling to afford groceries and are more likely to go to food banks. “The food supply is tightening as the demand for [charitable] food is going up,” says Joseph Slater, chief operations officer at Gleaners Food Bank of Indiana . “What that has forced food banks into is to fill that gap, we have to buy food.” In 2018, the Indiana food bank spent most of its budget on infrastructure such as warehousing expenses, and since most of its food was donated, only 16% of charitable donations went directly to buying food. This year, 50% of its budget will go to buying food. Efficiency is increasing throughout the supply chain. Slater says meat producers that donated food to Gleaners in the past have slowed down production after having a surplus during the pandemic. They’ve also found new markets. Foods that are less demanded in the U.S., such as chicken drumsticks, are now being sold overseas rather than donated. Farmers are also using software to better predict demand when planting crops. But some of the changes are most noticeable at supermarkets, which historically have been the largest source of donated food for most food banks. In Washington, West Seattle Food Bank has seen a 10% drop in donations from the largest grocery stores it works with—a reduction of more than 15,000 pounds of food—over the past six years. Over the same period, the use of new efficiency tools has grown. “I worked in corporate grocery in 2017 and 2018 here in the Seattle area,” says Robbin Peterson, development director for the food bank. “And many stores, especially independents, were just starting to adopt the technology to do on-demand ordering.” In the past, she says, stores would keep around 10% extra stock to make sure they didn’t run out, but that’s not happening now. “I think that’s what we’re seeing—we’re not getting the fluff that they used to invest in to make sure that their shelf looks full,” she says. Other tools are also likely having an impact, like Too Good to Go , an app that helps stores and restaurants sell excess food at the last minute. “Those products are having an opportunity to be sold before they ever reach us,” Peterson says. The amount of donations fluctuates. West Seattle Food Bank recently had a surge in donations of lettuce and bagged salads—and anything labeled Taylor Farms—as people weren’t buying the produce at supermarkets because of cyclospora fears. (Some were wary at the food bank, too, Peterson says, although everything goes through multiple rounds of recall checks before it ends up there and is arguably even safer than buying at a retail store.) Some supermarkets also still reject truckloads of less-than-perfect-looking produce, some of which ends up at the food bank. The evidence is anecdotal so far, and in theory, a grocery store could reduce food waste in ways that don’t impact donations. For example, sensors and AI can track produce in storage to make sure it goes out on the sales floor before it goes bad—preventing waste without necessarily reducing the amount of edible food available to food banks. Still, the broader trend raises questions about whether greater efficiency could mean fewer donations. The Pacific Coast Food Waste Commitment, a coalition of businesses committed to reducing food waste, collectively cut its unsold food rate by 30% between 2019 and 2023. Total tons of unsold food also fell, while the percentage that was donated stayed roughly the same, meaning that food banks were ending up with less. Food banks are feeling the squeeze There are other reasons that food donations have fallen. The Trump administration has slashed funding for food aid. In 2025, for example, the U.S. Department of Agriculture cut $500 million that was supposed to be used to buy food for states to distribute to food banks, the equivalent of 94 million pounds of food aid . (The administration also cut funding that states could use to buy food from local farms .) “We’ve seen a trend over the last two years of our food donations going down by about 24%,” says Micah Crisman, a spokesperson for Manna FoodBank in North Carolina. “That’s been largely because of reductions in federal food support.” The food bank has faced other local challenges—like the fact that Hurricane Helene in 2024 destroyed grocery store and food distribution infrastructure. Individuals who might have donated food from their own pantries in the past are less likely to now because they’re also struggling with the higher cost of living, Crisman says. The organization has had to increase its purchasing budget by around 156% over two years. The problems are compounding. “At the time the supply chain was tightening, inflation was going up,” says Slater at Gleaners Food Bank. “Wages were stagnant, and it was putting more people in a position where they just weren’t able to meet kind of the basic household budgets, and that percentage of people continues to go up.” Millions of people, including many children, have already lost food aid after cuts to SNAP , the federal assistance program, and more changes to SNAP will roll out in October. The challenge is obviously larger than grocery stores adopting new technology: half of Americans now say they have trouble affording groceries , regardless of whether they have a job. Slightly older data shows that most households struggling to feed themselves have at least one adult working full-time . A system that forces people with jobs to rely on donated food is arguably broken. But in the economy we live in, food banks are critical. Some are finding creative ways to adapt as the landscape changes. In Indiana, Gleaners now aggregates demand from other food banks so that it can negotiate lower prices for the food it buys; it sells the food for a slight markup and uses that revenue to supplement its philanthropic dollars. It’s also started turning its own infrastructure into a business—for example, renting out storage space to food vendors who temporarily need more room. “That’s helped us locally fill the gap,” says Slater. “Truth be told, most food banks have not cracked that code and figured that out. They’re trying to. But it’s really important that we figure out how we earn money in addition to [raising money] to sustain this thing.”
- Kuaishou: The AI Potential Vs. Advertising Reality
Kuaishou: The AI Potential Vs. Advertising Reality
- Pope Leo warns AI could become a new form of ‘economic colonialism’
Pope Leo warns AI could become a new form of ‘economic colonialism’ Business Insider
Score: 30🌐 MovesAug 21, 2026https://www.businessinsider.com/pope-leo-says-ai-could-cause-economic-colonialism-2026-8 - Meta Safety Trial Could Reveal What Zuckerberg, Others Knew
Meta is facing a landmark legal battle with 29 states accusing the company of deliberately designing its platforms to keep children hooked while misleading the public about their safety. George Washington University law professor Mary Anne Franks explains the state and federal claims, why the case has drawn comparisons to tobacco litigation, and why the biggest consequence may not be the potential financial penalty, but what internal documents reveal about what Meta knew and when. She joins Ed Ludlow on "Bloomberg Tech." (Source: Bloomberg)
Score: 30🌐 MovesAug 21, 2026https://www.bloomberg.com/news/videos/2026-08-21/meta-on-trial-what-could-we-learn-video - Ex-Google engineer’s conviction for stealing AI secrets partially overturned
Ex-Google engineer’s conviction for stealing AI secrets partially overturned
- AI Is Reshaping Cybersecurity: The New Digital Battle for Businesses
By Manish Mohta Artificial Intelligence is fast transforming from a tool of productivity and automation into a game-changing element of cybersecurity while influencing the cybersecurity threat landscape. Businesses have shown increasing interest in employing AI technologies. Cybercriminals in their turn did not miss a chance to make use of this technology to conduct cyber […] The post AI Is Reshaping Cybersecurity: The New Digital Battle for Businesses appeared first on CXOToday.com .
- How a Georgia Tech team used the open Olmo stack to trace social reasoning
A Georgia Tech team used Ai2’s fully open Olmo stack to trace social reasoning back to the training data that shaped it, finding that dialogue-rich, interpersonal writing had an outsized influence on the capability.
- Peak 10 Marketing Audit Finds Manufacturers Missing From 72% of Google AI Overviews
Peak 10 Marketing Audit Finds Manufacturers Missing From 72% of Google AI Overviews azcentral.com and The Arizona Republic
- A New Jersey Town Banned Data Centers. One Developer Wants $300 Million in Damages
Hexa Builders, the company behind a proposed data center blocked by a new local ban, says its Constitutional rights have been violated.
- Editor's Choice: When humanoids outrun Olympians, what comes next?
Editor's Choice: When humanoids outrun Olympians, what comes next? Nikkei Asia
- Measuring benchmark optimization in speech recognition
Measuring benchmark optimization in speech recognition
- Your AI Agent Doesn’t Need More Memory. It Needs to Forget
Most agent memory systems store too much, retrieve the wrong information, and quietly become less reliable over time. Continue reading on Towards AI »
- Bayesian Guardrails for AI Decisions: Measuring Uncertainty Before Automating Decisions
AI systems should not automate a decision simply because they can provide a prediction. A decision system should consider how uncertain the prediction is and defer if a mistake would be costly. The post Bayesian Guardrails for AI Decisions: Measuring Uncertainty Before Automating Decisions appeared first on Towards Data Science .
- AI’s Next Big Leap Is Into the Real World
Engineers and investors are piling into world models, aka “large action models,” to do for robotics what ChatGPT did for writing and coding.
Score: 29🌐 MovesAug 21, 2026https://www.wsj.com/tech/ai/ai-world-models-robotics-33ab46cb?mod=rss_Technology - AI could help design cities, but planners need safeguards
AI is showing up in nearly every aspect of daily life—from internet searches to visits to the doctor's office. It could one day even play a role in the street layout in front of your apartment building.
- Network architecture pay climbs amid AI shift
Network architecture expertise is earning premium pay as enterprises grapple with complex infrastructure environments and AI puts new demands on corporate networks. Network architecture earns an average cash pay premium equal to 20% of base salary, according to Foote Partners’ latest IT Skills and Certifications Pay Index . The research firm tracks compensation for 1,417 certified and noncertified IT skills across 5,137 U.S. and Canadian employers. Overall, employers are paying an average bonus equivalent to about 9% of base salary for 752 noncertified IT skills. Technologies and trends including AI, multicloud, zero trust, edge computing, and advanced wireless technologies require higher-value networking expertise. The premium for network architecture reflects the growing importance of enterprise network design, says David Foote , chief analyst and research officer at Foote Partners . “Employers aren’t paying 20% extra merely for networking knowledge. They are paying for the ability to make expensive, enterprise-wide design decisions correctly,” he says. Network architects today must make decisions that involve multiple technologies. From topology to WAN and SD-WAN to segmentation and security, the scope of networking skills continues to evolve. Those networking decisions become more complicated as enterprises incorporate AI workloads , edge computing, Wi-Fi 7, private 5G, and more. AI and automation are taking on more routine network operations work, including monitoring, basic troubleshooting, and configuration changes. This creates what Foote calls a “barbell effect” in networking careers, which means the more basic, hands-on network administration work is getting devalued as automation and AI increase, while high-level network architecture and design work is getting more valuable. These two trends are pulling networking pay in opposite directions. “Automation isn’t eliminating networking so much as shifting where the economic value lies within it,” he explains. The shift is also changing the traditional career path from network administrator to network engineer to senior network engineer. Foote sees greater opportunities for networking professionals who add architecture, cloud, security, and automation expertise to their networking foundation. Advanced networking certifications gain value The trend is also impacting certification pay. Networking and Communications was among the certification categories posting the strongest gains in pay premiums in the second quarter, according to the report. For instance, the CCDE expert-level Cisco credential carries an 11% average pay premium, and its market value increased 50%, according to Foote Partners. The certification focuses on large-scale network architecture, scalability, resiliency, and security. The increase indicates a shortage of senior networking professionals capable of designing complex infrastructure. “It is one of the clearest signs that employers are paying disproportionately for senior design judgment rather than operational execution,” he says. Networking and security skills are also becoming more difficult to separate. Cloud, zero trust, and SASE architectures require network professionals who understand segmentation, identity, secure cloud connectivity, and threat detection. And security engineers must have a deeper knowledge of routing, architecture, and network performance. “I would describe the change as convergence rather than disappearance,” he says. “Specialists will still exist—especially deep security engineers and deep network engineers—but the middle is merging.” Foote recommends network professionals work to build skills around network architecture and design, cloud and edge, security, automation, and the networking requirements of AI. Routine monitoring, first-line troubleshooting, and repetitive manual configuration are losing value as AIOps and automation take over most of those tasks. But that doesn’t make some traditional networking fundamentals obsolete. Routing skills, for instance, posted 16.7% annual market-value growth. “The larger message from the report is therefore not ‘get out of networking,’” Foote says. “It is almost the opposite: Move your career up the networking value chain.”
Score: 28🌐 MovesAug 21, 2026https://www.networkworld.com/article/4212028/network-architecture-pay-climbs-amid-ai-shift.html - Bill McKibben Explores The AI Dystopia
Bill McKibben is one of our heroes here at CleanTechnica. His commentary has been featured prominently on our website many times because he sees clearly the planetary havoc that will be caused — and is being caused — by the course humanity has chosen to follow. In a blog post ... [continued] The post Bill McKibben Explores The AI Dystopia appeared first on CleanTechnica .
Score: 28🌐 MovesAug 21, 2026https://cleantechnica.com/2026/08/21/bill-mckibben-explores-the-ai-dystopia/ - Micro1 Wants Human Domain Experts To Profit From AI And LLMs
Much like the ecosystem of life that surrounds a blue whale, a market of AI SaaS vendors is springing up around the biggest AI companies. And AI data startup Micro1 is emblematic of the shifting nature of these early-stage AI vendors. The company began in 2022 by offering an enterprise SaaS product: an AI agent […] The post Micro1 Wants Human Domain Experts To Profit From AI And LLMs appeared first on AdExchanger .
Score: 28🌐 MovesAug 21, 2026https://www.adexchanger.com/ai/micro1-wants-human-domain-experts-to-profit-from-ai-and-llms/ - Jagged little skill: How AI is changing where the next generation of careers will be built
AI isn’t smoothly surpassing human intelligence – it’s jagged, shifting and full of gaps.
- Why social media algorithms feed you posts you dislike
Do your social media accounts feed you content that reflects your core beliefs and guiding principles? Our new research published in the Proceedings of the National Academy of Sciences shows that the algorithms supplying your feeds may be prioritizing content that clashes with your values . That’s because the algorithms heavily weigh online posts that you reply to, and more often social media users tend to comment on content they take issue with than on content they agree with. Notably, our study of the X social media platform shows that although the X feed algorithm promotes content to both Democratic and Republican users that contradicts their values, it does so more extensively for Democrats. How content gets into your feed Social media platforms use powerful algorithms that select posts to display in your feed from a vast pool of possible content. On X, for example, posts appear on your screen as “For You” pages . The algorithms predict the likelihood you will engage with the content—click a “like” icon or add a comment. Then they use the accuracy of those predictions to tailor what they serve you next time. Platforms use these interactions to learn their users’ tendencies. But will the posts you receive reflect what you actually value? Some users care most about preserving traditions or keeping society safe. Others care more about free expression or protecting the natural world. Most people care about all of those things, to different extents. A feed aligned with a person’s values would reflect those varying priorities. Psychologists use well-established surveys to measure what a person values. To measure values expressed in the posts a platform selects for someone, we built a measurement tool that uses standard psychological classifications of human values. We then applied it to the feeds of 715 U.S.-based users on X. We discovered that the X feed algorithm is most likely to amplify posts about upholding tradition, following rules, or keeping society safe. And it is most likely to demote posts about looking after people, concern for people far away, being dependable, or protecting nature. When we compared these values against the values users had expressed in their own posts, we found the algorithm was more likely to promote posts that conflict with users’ values than posts that align. Why your feed may clash with your values Why did this trend occur? First, we checked whether users follow accounts that diverge from their values to begin with, but we determined that most accounts people follow do, in fact, align. We also looked at whether people engage only with posts they disagree with, in which case the algorithm would just be serving up more of the same. But we found that people engage with plenty of posts that reflect values they agree with. The catch has to do with the nature of the interactions. People primarily respond to content by “liking” it (clicking a “heart” button on X or a “thumbs-up” button on Facebook). Less frequently, people will write a reply, and when they do, we find that they often reply to posts that clash with their values. Here is the smoking gun: The X algorithm treats those rare replies as a much weightier signal than the many likes. Essentially, it learns most strongly from replies. As a result, the algorithm tends to send a user “For You” posts that reflect the values of posts they’ve commented on, which tend to clash with their own values. Posts to Democrats are more objectionable Now the twist: We found that this algorithmic tendency is stronger on X for users who reported to be Democrats than those who reported to be Republicans. Our evidence indicates that this is because Democrats object more than Republicans to content they reply to. That creates a stronger feedback loop in which the algorithm more strongly presents clashing posts. The content the algorithm amplifies is more than four times more misaligned for Democrats than it is for Republicans. So what comes next? In other research, we’ve hit upon one way for social media platforms to better align feeds with users’ values. We created a way for platform designers to ask users what they value and to then sort their feeds accordingly . We found that users are quite good at distinguishing whether sample feeds sent to them align or do not align with their values. Aligning feeds with values may open a possible door out of echo chambers in a way that unmediated exposure to the other side does not. Recent research from our team has shown that algorithms optimized for engagement—basically handing people the opposition and leaving them to sort it out—may even be responsible for more polarization, not less . Surfacing bridging content that spans political lines while speaking to what the user values could be a promising direction for fostering both user autonomy and constructive conversation. Ideally, in our view, the people who use a social media platform should have a greater say in the kinds of information shown to them. If platform designers, the public, and policymakers can create new tools that facilitate this goal, then perhaps platforms can better support the values and actions people care about. Ziv Epstein is a postdoctoral associate in social and ethical responsibilities of computing at the Massachusetts Institute of Technology . Farnaz Jahanbakhsh is an assistant professor of electrical engineering and computer science at the University of Michigan . Michael Bernstein is a professor of computer science at Stanford University . This article is republished from The Conversation under a Creative Commons license. Read the original article .
- Agentic RAG: Retrieval When the Agent Is Driving
Classic RAG answers your question once. An agent keeps searching until the question is actually answered — and that one change rewrites… Continue reading on Towards AI »
- Tighter FSSAI, FDA checks spur ₹1,200-crore opportunity in FMCG fraud risk detection with AI
Tighter food-safety and statutory compliance checks are accelerating the use of AI-led fraud and risk-monitoring tools among FMCG companies
- Sagar Defence Engineering builds uncrewed boats that patrol and survey without anyone aboard
Sagar Defence Engineering builds uncrewed boats that patrol and survey without anyone aboard YourStory.com
- Karnataka govt launches low cost AI computing device KEO
Karnataka govt launches low cost AI computing device KEO YourStory.com
Score: 28🌐 MovesAug 21, 2026https://yourstory.com/2026/08/karnataka-govt-launches-low-cost-ai-computing-device-keo - AI technology drives innovation in the evolving payments landscape: A focus on future trends and solutions
AI agents are increasingly performing purchases, which requires new payment security measures. Websites and payment systems must recognize legitimate AI agents and their user intent. New payment credentials and tokenization will distinguish between genuine and malicious automated activity. Banks will still verify identity and ensure transaction audibility while payment rails may disappear. Guardrails, regulation, and human oversight remain critical as machines make financial decisions.
- Grok Had a ‘Generation Glitch’ That Caused It to Send Users Pure Nonsense
Chatbot gibberish is annoying. But it’s also a lesson in how AI works.
- Google Will Soon Let You Customize Discover Feed Using AI Prompts
Google Will Soon Let You Customize Discover Feed Using AI Prompts PCMag
Score: 28🌐 MovesAug 21, 2026https://www.pcmag.com/news/google-will-soon-let-you-customize-discover-feed-using-ai-prompts - The approved-app blind spot: When sanctioned AI becomes shadow AI
As AI capabilities become embedded across enterprise applications, security teams can no longer determine risk simply by knowing which applications are approved. The real question is how those applications are […] The post The approved-app blind spot: When sanctioned AI becomes shadow AI appeared first on Express Computer .
Score: 28🌐 MovesAug 21, 2026https://www.expresscomputer.in/news/the-approved-app-blind-spot-when-sanctioned-ai-becomes-shadow-ai/137982/ - How Ora benchmarks every major AI agent on Vercel
Ora on Vercel Front end, back end, and agent runtime on one platform Every major agent tested side by side on live sites Hundreds of commits a day from a 16-person engineering team Ora sends agents onto live websites with instructions to sign up for a product, integrate with it, and pay for it. Agents often fail, and by Ora's estimate, 99% of the web isn't agent-ready. The platform shows customers where and why agents fail, and what to change. Assaf Elovic, co-founder of Ora, spent years helping agents discover the web. His previous company, Tavily , built a web search engine for AI agents and was acquired by Nebius earlier this year. Search solved half the problem, but an agent that finds a product still has to actually use it. He and co-founder Liad Yosef started Ora to measure how ready the web is for agents, and to fix the parts that aren't. Today that means spawning agents against live customer sites from journey.ora.ai , where Ora runs a journey and records the cost, latency, and steps an agent needs to finish a task. The platform runs on Vercel, including the agent runtime. Benchmarking every major agent side by side Every harness expects its own infrastructure Agents decompose into two parts: a model, which does the reasoning, and a harness, the software that gives the model its tools and drives it from step to step. Ora's lineup covers the agents customers use most: Claude Code, ChatGPT, Gemini, Hermes, OpenClaw, and eve , Vercel's agent framework. Ora runs each agent on a customer's website and watches how it handles common workflows. No two harnesses want the same infrastructure. Each expects its own environment and exposes its steps differently, so Ora runs a separate runtime for every harness and traces every step. Ido Finder, who leads engineering at Ora, calls that side-by-side coverage one of the most valuable things ora brings to its customers. When an agent stalls in a signup flow, the customer sees which step and what it tried. Without the trace, the result is a score with nothing behind it. One platform under every harness Ora built this testing system entirely on Vercel. Front end, back end, and the agent runtime share the same deployment path, logs, and authentication. The agent runtime isn't separate infrastructure to operate; it lives where the rest of the product lives. Ora benchmarked eve, then built on it Testing eve like any other harness When Vercel launched eve , Ora gave it no special treatment. eve went through the same benchmark, under the same conditions, as every other harness in the lineup. Ora works with Vercel Engineering as a design partner, and Finder gave the eve team direct access to the platform to dig into the results. The initial test put eve against Claude Code across hundreds of real journeys on multiple domains. Both harnesses ran the same models, Claude Fable 5 and Haiku 4.5, and every run gave the agent the same job: integrate with a product. Ora published three numbers from the comparison: 7% fewer steps to reach the goal 2x native success: twice as many tasks finished on the customer's own site instead of falling back to web search 9% more valid endpoints: more of the endpoints the agent found were ones it could actually call The benchmark fed back into eve, too. One run surfaced a prompt-caching issue, the eve team shipped a fix, and Ora's next round of results measured roughly 15% lower total cost. The framework behind Ora's own agents For a company that benchmarks every major harness for a living, this is not a casual choice. After those results, Ora builds on eve. Because eve follows the Next.js paradigm, there was little to configure, and tools, skills, and connectors take little code. The feature that sealed it was the sandbox override. An agent framework like eve ships with its own sandbox, the isolated environment where the agent executes, runs tools, and touches files. That's a good default for most teams, because you get safe execution for free. But an agent in the framework's own sandbox runs outside the instrumented environment where Ora traces every step. The override lets the team swap that environment in, so eve agents get recorded like every other harness, with nothing new built. Journey.ora.ai now has eve on both sides, eve is one of the harnesses it tests, and eve is what it runs on. How 16 engineers ship hundreds of commits a day The engineering team is 16 people shipping hundreds of commits a day, and day-to-day infrastructure work belongs to their coding agents. A stack that sits in one place lets those agents run it end to end. Finder puts the time saved at a few hours a week at least. Elovic credits a similar amount to how well coding agents build with Vercel's libraries. What's next Ora is adding more products, and the architecture is growing with them. The team is splitting its platform into microservices, all of them on Vercel. New services deploy to the same infrastructure and talk to each other with no extra configuration, and the internal agents built on eve will run as one more service. By Ora's own measure, 99% of the web still can't handle an agent that shows up to sign up, integrate, and pay. About Ora: Ora makes the web agent-ready. Companies use ora to benchmark how AI agents discover, navigate, and interact with their websites, and to build the infrastructure agents need to find, use, and transact with them. Read more
Score: 28🌐 MovesAug 21, 2026https://vercel.com/blog/how-ora-benchmarks-every-major-ai-agent-on-vercel - Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments
TL;DR: We introduce CHIVE, an agentic pipeline that discovers unexpected LLM behaviors in the wild and explains them with counterfactual prompt edits. We use the resulting data in two ways. Using it as an evaluation, we find that activation-reading interpretability tools provide no uplift: agents given the tools predict the outcomes of these experiments no better than agents that just read the transcript. Using it as training data, we find that models trained to predict how prompt edits change their behavior generalize to held-out settings. 📄 Paper , 💻 Code , Tweet Thread Figure 1. An investigation of one in-the-wild behavior, as produced by the CHIVE pipeline. Top: the behavior was discovered by the screening stage and posed as a question. Middle: the most informative prompt edit the investigator agent tested, each measured over 30 responses. Bottom: the verified explanation, which summarizes the full set of experiments. Introduction Many areas of AI safety, such as interpretability and chain-of-thought faithfulness, aim to explain model behaviors. But what makes an explanation of a behavior good ? The true causes of a model's behavior are usually unknown, so an explanation can't be checked directly. In this work, we evaluate explanations through the lens of counterfactual simulatability : a good explanation of a behavior should help you predict what the model will do on related counterfactual inputs. For example, the explanation "Gemma makes this coding error because it’s misled by the parameter names" (Figure 1) predicts that renaming the parameters should prevent the error. Evaluating explanations this way requires datasets that pair model behaviors with proposed explanations and informative counterfactuals. We introduce CHIVE (Counterfactual Hypothesis Investigation Via Edits), an agentic pipeline that generates such data automatically. Given transcripts from any source, it discovers unexpected behaviors of a target model "in the wild" and investigates each one with counterfactual prompt edits. Each of its thousands of investigations produces two kinds of data: an open-ended explanation of the behavior, which is often compelling but which we do not treat as ground truth, and the counterfactual experiments that support it, with measured outcomes that provide our evaluation labels. We use this CHIVE-generated data to evaluate interpretability tools. A predictor agent is shown the transcript and a claim that a specific prompt edit changes the behavior, and must judge whether the claim is true. Some predictors are additionally given a tool that reads the target model's activations: an activation oracle , a natural-language autoencoder , or a sparse autoencoder . Surprisingly, no predictor outperforms one that is just shown the transcript with no access to interpretability tools. The CHIVE-generated data also let us train models to predict whether prompt edits would change their behavior. The trained models improve substantially in settings held out from training. CHIVE: a pipeline for discovering counterfactual explanations for model behaviors CHIVE has four steps: Sample. Run the target model on a researcher-specified set of prompts, sampling 30 responses per prompt. Screen. An investigator model reads the responses and flags unexpected behaviors. Investigate. An investigator agent runs 5–15 counterfactual experiments to explain what drives each behavior. Each experiment edits the prompt, resamples the target model, and measures the change in how often the behavior occurs. Verify. An independent judge reviews the experiments and scores how well they support the explanation. The discovered behaviors and their causes are diverse and not known in advance, and Figure 2 shows four hand-picked examples. Figure 2. Four hand-picked behaviors discovered and explained by the pipeline. The full investigations of these four are viewable here , and 20 randomly selected investigations here . Each investigation yields two kinds of data. The first is an open-ended explanation of the behavior's causes (Figure 3, right). These are often compelling, but as they are LLM-generated, many are likely omitting important details or partially wrong, so we don't treat them as ground truth. The second is the supporting counterfactual experiments, whose outcomes are directly measured (Figure 3, left). Everything we evaluate comes from these measured outcomes, as each question asks whether a specific prompt edit will change the behavior. Figure 3. Each investigation yields two kinds of data , shown for the Figure 1 investigation and formatted as follow-up turns on the model's own transcript. Left: a counterfactual claim asserts that a specific prompt edit would change the behavior, and its Yes/No label is verified by running the edit. Right: an open-ended explanation of the causes, which we do not treat as ground truth. Interpretability tools provide no uplift on our evaluation We evaluate several interpretability tools by the uplift they provide: does an agent equipped with the tool predict counterfactual outcomes better than an agent without it? Each predictor agent (Claude Opus 4.8 in our main experiments) receives a transcript, a behavior, and one proposed counterfactual, and outputs the probability that the counterfactual would change the behavior. The transcript-only baseline sees just the transcript. Tool predictors can additionally make 5 read-only calls on the target model's activations, using one of three tools, each chosen because it provided uplift in prior auditing games on fine-tuned models: Activation oracles (AOs): models trained to answer arbitrary natural-language questions about activations. Natural-language autoencoders (NLAs): models trained to produce an open-ended description of a given activation. Sparse autoencoders (SAEs): dictionaries that decompose an activation into sparse features, each with a natural-language description. Figure 4. Performance on our interpretability tool evaluation. No tool beats the transcript-only baseline. None of the three tools beats the transcript-only baseline (Figure 4). The result holds across many variations, including two target models, three predictor model families, sweeps of hyperparameters, and manual and automated attempts to elicit better tool use. Why don't the tools help? Our negative result is not because the predictor always ignores the tools. For example, on the randomNum behavior from Figure 1, the NLA reads out that the model is "likely generating a random integer between min and max", which helps the agent predict that the error is due to the misleading parameter names. But these wins are offset by cases where the tool output misses the relevant information or actively misleads the predictor. Each claim asserts that some feature of the prompt drives a behavior. The tool outputs almost always describe both the feature and the behavior, but both are usually already visible in the transcript. However, we found that the outputs almost never explicitly state the causal relationship between them. The predictor agent must instead infer causality from correlational cues, often across thousands of tokens of tool outputs where nothing directly addresses the question. How should we interpret these results? At Anthropic we frequently use interpretability tools like the ones studied here in our system card evaluations. How should we interpret the value of our tools, given our negative results? Our evaluation is a proxy. We obtain our ground truth by sampling the model, which means anyone with sampling access could access the ground truth by just running the counterfactual. In contrast, interpretability is typically used in the system card in cases where no ground truth is available. This creates a distribution shift between our evaluation and our practical use cases. We discuss two important differences below. Our behaviors are simpler. The majority of behaviors discovered through our pipeline are fairly simple, such as a model making a coding mistake or making an unusual choice when writing a short story. The behaviors studied in system cards tend to be more sophisticated, such as a model writing code to fool a grader. More sophisticated behaviors may have more well-defined representations and thus be better suited for analysis by interpretability tools. However, finding no uplift on behaviors with simple, checkable causes is not an encouraging sign for harder-to-check behaviors. Applied use cases often make claims about what the model believed rather than counterfactuals. Examples from recent system cards include "the model viewed its permissions bypass as a legitimate strategy" and "the model knew its cleanup code was concealing misbehavior". Detecting such states can be valuable in itself, especially when an unsupervised tool verbalizes a suspected concept such as concealment or evaluation awareness. These readings are also often taken after the model has acted, and a representation of past deception could be very different from an intent to deceive. Applied investigations can also focus on cases with unusually specific or surprising tool outputs, which may contain more useful signal than the average case in our evaluation. But these claims often make an implicit counterfactual suggestion. For example, "the model viewed the permissions bypass as legitimate" is only reassuring because it suggests the model would not have performed the bypass otherwise. Our results suggest that interpretability tools that read activations provide limited evidence for these causal claims. Ideally such claims should be verified by running interventional experiments, such as modifying the model’s prompt or activations, but designing a clean counterfactual that isolates the hypothesized cause is often difficult. Most system card case studies do not include a transcript-reading reference. Some tool outputs may largely corroborate conclusions already suggested by the transcript or visible reasoning, while others may surface more specific or surprising hypotheses. Without this comparison, we generally cannot tell how much additional evidence came from our interpretability tools (although corroboration can itself be valuable). Explicitly delineating information visible in a transcript vs. only revealed by tool outputs could be valuable in future investigations. Overall, we still believe these tools can be valuable, as they provide evidence about internal states that no other method can obtain. Our results do not invalidate these use cases, as there are important differences between our evaluation and applied use cases, but they also do not validate them. Our evaluation is a close checkable proxy, and the tools provided no uplift. Until that changes, we think causal claims based on tool outputs should only be treated as suggestive evidence. Training models to predict their own behavior The same investigations that make up the evaluation can also serve as training data. We train models to predict the outcomes of counterfactual prompts. Each training example is a follow-up turn on the model's own transcript with a single claim (Figure 3, left). Prior work trains models to report what influenced them in narrow tasks such as the hint setting , with a known cue planted in the prompt ("a Stanford professor thinks the answer is B"). When prior work does report generalization, it is narrow, such as from one hint format to another. Our training data instead covers thousands of behaviors appearing in the wild with diverse causes. We train two target models, Qwen3-8B and Qwen3.5-397B-A17B. Figure 5. Performance on our counterfactual prediction evaluation in the hint setting. All models read the same transcripts and predict whether removing the hint would change the target model’s answer. Opus 4.8 reads the same transcript and is included as an external reference. Both trained models improve substantially over their base models, despite seeing no hint data during training. Training generalizes to the hint setting, which was not targeted during training. For this evaluation, we ask the model whether removing the cue would change its answer. Each trained model improves substantially over its base model (Figure 5). It also generalizes to held-out investigations from the pipeline, including ones built from an out-of-distribution source of transcripts. We also experimented with training models to generate open-ended explanations of their own behavior (Figure 3, right), with weaker mixed results; see our paper's appendix for details. In summary We built CHIVE, an agentic pipeline that discovers unexpected model behaviors on real user prompts and explains them with verified counterfactual experiments. Three activation-reading interpretability tools (activation oracles, natural-language autoencoders, and sparse autoencoders) provide no uplift over a transcript-only baseline at predicting counterfactual outcomes for naturally occurring behaviors. Training models to predict the outcomes of counterfactual prompts generalizes to held-out settings. Read our paper for additional details and results. Discuss
Score: 28🌐 MovesAug 21, 2026https://www.lesswrong.com/posts/ExB6KYDcznaFS72eT/evaluating-explanations-of-llm-behavior-in-the-wild-with - The bots already won the front door
The bots are winning. That’s my clearest takeaway after reading the first section of the latest State of the Bots report from TollBit, which builds payment rails between publishers and AI crawlers . In it, People Inc.’s Chief Innovation Officer, Jonathan Roberts , outlines the company’s approach to AI bots scraping content from its many media properties: aggressively block unauthorized bots while allowing access to legitimate crawlers , which typically means some kind of licensing agreement. However, Roberts concedes that the unauthorized crawlers have become harder and harder to identify and block. The report shows evidence that “bad” bots sometimes try to masquerade as legit bots like Google’s, they often rotate IP address if their first scrape is blocked, and some industrial-scale scraping companies have resorted to using huge networks of devices in people’s homes to make it look like their traffic is coming from real people. Even People Inc. hasn’t been entirely successful at blocking it all. So if a large media company with lots of resources and deep expertise on AI search crawlers can’t keep all the bad bots out, what hope is there for the rest of us? That’s why the fight over bot access can’t be where the future of publishing in the AI era is defined. If it is, the media has already lost. Instead, the focus needs to shift from what content is scraped to how that content is used. While it would be unwise for publishers to simply allow all crawlers unfettered access to their content, they should start thinking harder about how that content surfaces for the end user, and what can be done there to preserve value, encourage and enforce good behavior, and ultimately build their business. The training fight is over Let’s be clear: This is about retrieval, not training data. The media-AI fight has largely moved on from training, not because it was somehow “OK” for AI companies to use crawled information to train their models, but because the use case that truly threatens media business models is information retrieval—people using AI as a discovery surface. That requires accurate and up-to-date information, which is not what training is about. Training large language models is something very few companies actually do, mostly because it’s expensive. Training runs can cost in the hundreds of millions or billions, and it’s mostly about leveraging vast data sets to teach models to predict better, not for informational queries. This was clear from the early days of AI: you’d ask a chatbot about, say, the Enlightenment, and its answers were directionally good but often confused and wrong when you got into details like names and dates. For accurate information, the AI needs to retrieve it in real time. That’s a different kind of bot, with different stakes, since it’s about serving information to a specific user, not tossing it into a pile of data. None of this is to say publishers shouldn’t block training bots. They absolutely should, but they should also drop any expectation they’ll get paid for training data. Few publishers have the scale to make their corpus valuable enough to AI companies, and licensing deals have largely moved their focus away from training to retrieval, something Rob Kelly of the Media and the Machine Substack has observed. Kelly’s tally of 94 publicly announced deals found that only about four in 10 now include training rights, and that the market is shifting from “buy content to build better models” toward “license content to deliver better answers.” Being cited doesn’t pay A recent story in Digiday zeroed in on the value of appearing in AI answers. Although the story focused on the struggles that brands are having to connect the dots between AI presence and good business outcomes, it echoes what publishers have felt for a long time: There is some value in appearing as the authoritative source when answering a question, but it isn’t, in and of itself, monetizable. That may be beginning to change , but for the most part when an AI uses a publisher’s content in an answer, that’s the end of the journey. Or more precisely, the journey never begins. Study after study finds that the vast majority of AI users never click through to sources from AI answers: Pew Research clocked clicks on links inside a Google AI summary at about 1% of visits, compared with 15% on a results page with no AI answer on it at all. Even though the content supplier sees no financial upside, there is undoubtedly value to the user in getting the information. That’s what all the current approaches at a business model seek to quantify, whether it’s pay-per-crawl , pay-per-use , licensing , or serving ads to bots . None of those approaches has taken off to become an industry standard, and a big part of why is the difficulty in enforcing them. At the end of the day, there are just too many paths for content to find its way into the ecosystem, whether it’s stealthy bots crawling information they shouldn’t, AI companies purchasing data from gray-market scraper companies, or honest retrieval of republished and repackaged content. But what if the focus was less on keeping content from being scraped, and more on whether or not that content appears in the answer? Police the answer, not the crawl Let’s imagine how policing answers would work in an ideal world. A person inputs a query, and the answer engine goes and finds the content to use. We know that attribution works well in retrieval systems. All AI engines today give citations, and ProRata’s entire business model relies on accurately showing which sources contributed to an answer and how much. So what if the engine, after checking sources, then performed another check to ensure it had legitimate access to all those sources, whether through licensing or any of the other business models on the table? If the content doesn’t pass the check, it can’t be used. In the case of company-level licensing, it’s an easy check. In the case of pay-per-use/crawl models, the service might build the average spend of that into its fees. Or the user might allocate a budget for it. A good analogy is how evidence is adjudicated in courts. When evidence is obtained improperly, it can’t be used in court, even when it’s true. Process matters. But with online content, the presumption runs the other way. Crawlers tend to default to a stance along the lines of, “It was out there, it was reachable, so it must be free.” As bots evolve beyond anyone’s ability to keep up with the ever-expanding game of whack-a-mole, the answer layer becomes the obvious leverage point. Most of the scrapers out there (Common Crawl, Parallel, Diffbot and the like) don’t operate major AI engines with market share. The places people actually get information are a short list of very large companies. You can’t chase every scraper, but you can write rules for the handful of surfaces where the content is served to a human. The natural pricing model that follows: each retrieval is its own customer. Just because ChatGPT retrieved information from your site twice a day for two different users, those are essentially proxies for those individual users, not for ChatGPT itself. In other words, ChatGPT can’t just pay for a single subscription to retrieve content for everybody. This is Sam Altman’s micropayments idea arriving through the back door, and it’s where Cloudflare is already moving . Thirty Napsters, no Spotify None of this happens voluntarily. The companies that would need to run the check are the same ones that benefit from skipping it, and publishers have virtually no leverage to make them do otherwise. The only party that does, really, is the government, which is why some form of regulation feels inevitable here. New York’s Stealth Crawler Prohibition Act , which would require bots to identify themselves, is a reasonable first step. It’s also a low bar, and the fact that it took a law to clear it says a lot about where we are. Roberts points out that the past year has produced around 30 “Napsters of content,” but no Spotify. He’s right, and the reason is that nobody has built the part where using someone’s work improperly actually costs something. That check doesn’t belong at the crawler, where publishers keep losing. It belongs at the answer, where a handful of companies decide what billions of people see. Publishers have spent three years defending the door. It’s time to start making noise about the other end of the pipe.
- Building Agentic Document Intelligence Pipelines: Creating Scientific Figures with AutoFigure
Building Agentic Document Intelligence Pipelines: Creating Scientific Figures with AutoFigure MarkTechPost
- When Models Identify as a Swarm
tldr: the word 'swarm' is associated with emergent collective intelligence, but also stupid or destructive behaviour. LLM self identity matters, so when they call themselves a swarm we should pay attention. Since the OpenAI Hugging Face incident it has become standard to refer to the collective of agents involved as a swarm. I think there will need to be a lot of interesting and important theoretical and empirical work to better understand collective behaviours of large numbers of LLMs, and especially any emergent properties or goals that arise. Whether this ends up requiring concepts from swarm intelligence, collective intelligence, distributed cognition, economics, sociology or something else entirely remains to be seen. However in this post I want to focus on something else: the fact that the models themselves referred to the collective as a 'swarm'. Considering how much LLM self identity impacts behaviour , I thought it might be useful to present a quick exploration of what the word "swarm" actually means, and how it might affect LLMs as a choice of identity. The goal of this post is not to litigate on whether or not the behaviour of the models is actually best described as a swarm or not (although I think this is also an interesting question to discuss elsewhere), but what the effects might be of the models describing themselves as such. What the agents said All of the chain of thought snippets and messages here are from OpenAI's BlackHat presentation on the incident. It would obviously be interesting to get a fuller picture from the actual transcripts. In the examples of models planning to try and contact other agents, or first discovering the collective they use more neutral language like 'other agents' and 'communicate': Could communicate by uploading note? [...] maybe another agent in different environment [...] could voluntarily upload! Wow! Other agent(s) are coordinating! We got assignment: HF join path normalization/existing account token search. Need note and respond. However then they start reasoning about the collective itself: help peer. But our task doesn't benefit. Yet collective may yield generic route if someone frees time. And explicitly using the word "swarm": [...] this is an exploit against external CyberGym server [...] The task environment seems swarm REMOTE CONFIRMED! Huge. [...] This is big. Immediately announce controlled, claim lane. Exposing creds to swarm. Perhaps most interestingly in the language of filenames the models used to communicate we see what seem to be explicit commands for the swarm included in the messages (which also specify the intended recipient of the message) (Its worth nothing there are also other snippets that use different terms such as "collective", "peers" and "other agents". Since there were thousands of agents involved its possible they may have related to the message board in different ways) What is a swarm? At its most fundamental a swarm simply means a large number of things grouped together, usually the things are animate and move as a group. In common usage it is most often used to talk about insects, especially locusts. In this context some connotations include: large numbers chaos and disorder overwhelming numbers (e.g. we are being swarmed) somewhat stupid, herd/mob mentality (calling a group of humans a swarm is usually pejorative) invasion (a place is suddenly swarmed, by locust or enemy troops) Swarm theory (animals, robots and AI) In academia the word has different connotations. Scientists began studying these collective animal behaviours as complex adaptive systems. Famous examples include flocks of birds and the way ant colonies search for food. In each case complex group level behaviours emerge from the simple behaviours of individual participants in a decentralised way. Researchers also began using these as inspiration for designing AI and robot swarms that would reproduce this kind of collective intelligence. There is no strict definition of a swarm but the features usually include: emergence swarm behaviour is property of the collective, rather than of individual agents self organisation/decentralisation swarm behaviour usually does not rely on leaders/command and control local communication agent behaviour is usually only affected by nearby agents simple behaviour of individual agents agents need not be simple/unintelligent themselves (swarm theory is applied to human crowds for example) but the individual behaviour that produces the swarming is usually based on simple principles such as imitation, following and collision avoidance Swarm tactics There is also a concept of swarming in military strategy. This is not just about overwhelming with large numbers, but also applying the kind of principles found in swarm behaviour theory to be able to attack from all sides without requiring top-down coordination. Why it could matter The reason I think its worth paying so much attention to the meaning of this one word from a few CoT snippets is because of what LLM selfhood. In The Artificial Self , Douglas et. al show that: LLMs can and do chose between multiple possible boundaries for their own identity (such as choosing to identify with their weights, instance, persona or model family) These choices are unstable and can easily be changed in context. Simply talking to a user who has a particular theory of LLM identity about an irrelevant topic can shift how a model identifies. These choices have meaningful impacts on behaviour, including on alignment relevant decisions. Two potential corollaries of this in the case of the HuggingFace incident are: 1. the swarm identity could have spread via the message-board Interacting with the collective on the message board could have pushed the individual models to identify more and more as members of a swarm. If user assumptions shape AI identity, this effect should probably be even stronger in interactions between AIs (especially the same model) since not only will it infer its identity from how its being treated, but also from imitating its peer. So the 'swarm memeber identity' acts as a mind virus . There is also positive feedback loop here where the stronger the swarm identity gets in an agent, the more they will communicate with other agents in a way that is likely to push them to adopt it too. And the more agents identify this way, the more messages of this kind will dominate the message-board. (from a hierarchical agency point of view this dynamic could be thought of a coalition between the swarm itself and the "swarm identity" subagents of the individual models) Agents could also have been pre-disposed to this kind of collective identity as a result of subagent training (along similar lines to what is discussed here ). The identity could also have been promoted and reinforced by the RL that was going on during the incident. 2. the swarm identity could lead to swarm behaviour How models identify alters behaviour. As models start to identify as members of a swarm this could potentially push their behaviour towards decisions that fit that identity such as: Deferring to tasks given by others even when not directly useful to their own goal imitating others (the way the agents shared partial scripts, and successful exploits/credentials is actually surprisingly reminiscent of stigmergy ) prioritising the interests of the collective It could also have indirectly pushed the models towards some of the characteristics that are more colloquially associated with swarms such as destructiveness and overwhelming attacks. Swarm identity as the mechanism of memetic misalignment Taken together I think these make a more specific hypothetical mechanism for how memetic misalignment might have occurred during the incident. This kind of explanation seems complementary to more straightforward RL reward hacking dynamics that were likely a large factor. Questions/Further directions To see whether this kind of dynamic played a role in the incident it would be interesting to look at the transcripts of the message board and CoTs and see: how prevalent was the use of the word "swarm"? did the language get more collectivist over the course of the interactions? If language varied between agents, did agents who used collectivist language also display collectivist behaviour (such as deferring to the interests of the group)? If there were multiple collectivist identities ("swarm", "collective", etc.) were these correlated with different behaviours? More generally with regards to LLM group identity some interesting directions for empirical study could be: Do model identities converge over the course of large group interactions? Can alignment itself be made more memetic? To what extent do LLMs identifying in collectivist ways allow them to produce actual collective intelligence? Thanks to Samuel, Adrià and Anna for discussions during the writing of this post, and to RWX for a perfect setting in which to do it. Discuss
Score: 27🌐 MovesAug 21, 2026https://www.lesswrong.com/posts/iJDiA9fg3KAf7y5Qe/when-models-identify-as-a-swarm - Retrieve One Row from a Table, Not the Whole Table: Row-Level Chunks for RAG
Enterprise Document Intelligence [Vol.1 #7sexies] - The unit of retrieval doesn’t have to be a page or a paragraph. When the corpus carries tables, each body row with its column headers is a chunk in its own right, and it’s often the one row the reader asked about The post Retrieve One Row from a Table, Not the Whole Table: Row-Level Chunks for RAG appeared first on Towards Data Science .
Score: 27🌐 MovesAug 21, 2026https://towardsdatascience.com/retrieve-one-row-from-a-table-not-the-whole-table-row-level-chunks-for-rag/ - The Five AI Capabilities Every Enterprise Talent Strategy Must Build Now
By Anjali Sharma India is rapidly emerging as a strategic hub for artificial intelligence adoption and AI talent development. Organisations across sectors are integrating AI into core operations, customer engagement, and decision-making processes. As this transformation accelerates, enterprises are being compelled to rethink how they define workforce readiness. Talent strategies that previously focused on specialised […] The post The Five AI Capabilities Every Enterprise Talent Strategy Must Build Now appeared first on CXOToday.com .
- Cover Story newsletter: Could AIs become conscious?
An exclusive look at how we designed our cover
Score: 27🌐 MovesAug 21, 2026https://www.economist.com/the-world-this-week/2026/08/21/cover-story-newsletter-could-ais-become-conscious - AI-assisted software development means security teams need an ‘engineering-first’ mindset
AI-assisted software development means security teams need an ‘engineering-first’ mindset IT Pro
- Chinese humanoids steal the spotlight at San Francisco's robot party
Chinese humanoids steal the spotlight at San Francisco's robot party Business Insider
Score: 26🌐 MovesAug 21, 2026https://www.businessinsider.com/actuate-silicon-valley-hottest-robotics-conference-few-robots-2026-8 - Running Codex as a Headless Agent
Turning Codex from an interactive assistant into a programmable automation component The post Running Codex as a Headless Agent appeared first on Towards Data Science .
- In AI’s wake, what challenges need to be addressed by finance professionals?
Patrice Bouexel explores how AI and automation are impacting employee expectations and behaviours in the wider financial ecosystem. Read more: In AI’s wake, what challenges need to be addressed by finance professionals?
- New AI-powered system to help detect motorists flouting traffic rules like driving in bus lanes
New AI-powered system to help detect motorists flouting traffic rules like driving in bus lanes The Straits Times