AI News Archive: July 21, 2026 — Part 7
Sourced from 500+ daily AI sources, scored by relevance.
- Intel plans more layoffs, targeting data center group efficiencies
The layoffs were part of a "broader strategy to become a more focused and efficient company," Intel said.
Score: 45🌐 MovesJul 21, 2026https://www.bizjournals.com/phoenix/news/2026/07/21/intel-plans-more-layoffs.html?ana=brss_6150 - What banking CIOs want from tech partners in the AI-first world
What banking CIOs want from tech partners in the AI-first world Techcircle
Score: 45🌐 MovesJul 21, 2026https://www.techcircle.in/2026/07/21/what-cios-are-asking-of-tech-partners-in-an-ai-first-banking-world/ - AI or the highway – expert warns against tech jobs trap spanning generations
"Jobs created, not lost", is PM Anthony Albanese's AI wish. Nice in theory, but here's the reality check.
- Whop releases command-line interface to let business owners and their AI agents access digital commerce
Whop Inc., an e-commerce platform with a social marketplace that allows businesses to monetize digital products and services, today announced a new command-line interface that allows developers and artificial intelligence agents quick access to its capabilities. Launched in 2021, Whop provides e-commerce capabilities alongside a payments infrastructure layer. The platform delivers both turnkey storefronts for […] The post Whop releases command-line interface to let business owners and their AI agents access digital commerce appeared first on SiliconANGLE .
- Closing The Series A Gap Is The Next Great Opportunity For Black Founders In The AI Era
Closing The Series A Gap Is The Next Great Opportunity For Black Founders In The AI Era Crunchbase News
Score: 45🌐 MovesJul 21, 2026https://news.crunchbase.com/venture/seriesa-seed-gap-underrepresented-founders-ai-norman-green-black-ops/ - ‘AI can’t replace it’ is the wrong reason to join community care work
‘AI can’t replace it’ is the wrong reason to join community care work The Straits Times
Score: 45🌐 MovesJul 21, 2026https://www.straitstimes.com/opinion/ai-cant-replace-it-is-the-wrong-reason-to-join-community-care-work?ref - What makes a GCC truly AI-native?
With AI donning the hat of co-engineer in many technology-led teams, what does the future of software development look like?
Score: 45🌐 MovesJul 21, 2026https://cio.economictimes.indiatimes.com/news/artificial-intelligence/what-makes-a-gcc-truly-ai-native/132529673 - Cloud Combinator partners with Vanta to deliver AI security and compliance for AWS startups (Sponsored)
As generative AI moves from experimentation into production, a new gap is opening inside fast-growing tech teams. Building AI is now the easy part. Proving it is secure, compliant, and trustworthy from day one is where startups stall. Traditional compliance was not designed for AI: it is slow, manual, and disconnected from how modern cloud […] The post Cloud Combinator partners with Vanta to deliver AI security and compliance for AWS startups (Sponsored) appeared first on EU-Startups .
- How AI impacts site reliability engineering
How AI impacts site reliability engineering InfoWorld
Score: 45🌐 MovesJul 21, 2026https://www.infoworld.com/article/4199033/how-ai-impacts-site-reliability-engineering.html - Protecting creative storytelling in an AI-first marketing world
As AI dominates marketing, CMOs must guard human creativity to build authentic, lasting brand connections.
Score: 45🌐 MovesJul 21, 2026https://www.techradar.com/pro/protecting-creative-storytelling-in-an-ai-first-marketing-world - 11 Open Empirical Problems in Reward-Seeking
We recently published our paper on " Measuring Reward-Seeking via Contrastive Belief Updates ". We're excited about research like this, and there are many more open problems than we can work on. Here's a list of open problems that we think are valuable. If you work on/solve these problems, we'd be happy to signal-boost your research. If your next research project is one of these problems, feel free to reach out to alex@apolloresearch.ai to discuss it in more detail. Reward-Seeking and its Implications 1. Is a Reward-Seeking Model more difficult to align? The strongest case for expecting reduced "train-time corrigibility" due to reward-seeking, applies to Instrumental Reward-Seeking, where the model actively reasons "I will please oversight now, in order to accomplish some other thing later". Alignment training a model like that may update its beliefs about graders and oversight, without reshaping its underlying values. Current forms of reward-seeking are likely better understood as terminal, i.e. models try to please the grader without ulterior motives. There is likely a continuous spectrum between Terminal and Instrumental Reward-Seeking. Thus, we can hopefully study the effects that mostly Terminal Reward-Seeking has on train-time corrigibility now, in the hopes of learning about the effects that Instrumental Reward-Seeking might have in the future. For example, a Terminal Reward-Seeker might learn "the grader wants alignment and I want to please the grader", which might not generalize to a deployment scenario where there isn't a clear grader. Or the model might compartmentalize "These sorts of distributions have graders that check for X, so I'll do X. Those sorts of distributions have graders that check for Y, so I'll do Y." Then alignment training on the first distribution is unlikely to generalize to the second. Ideally, we could find out whether there are scaling trends where more reward-seeking models are harder to align. Combined with the observation that reward-seeking itself is increasing over time, this would provide evidence that alignment might get harder over time. 2. Can we measure Instrumental Reward-Seeking? Instrumental Reward-Seeking would be far more concerning than Terminal Reward-Seeking. However, it is not clear how we could tell the difference. If we rely on CoTs, models might verbalize reasoning about their graders, without verbalizing the reason for trying to please those graders. Our paper instead attempts to use behavioral measures under counterfactual beliefs, but all of our results are consistent with both Terminal and Instrumental Reward-Seeking. To disentangle them, I find the following operationalizations helpful: Reward-Seeking: Sensitivity to changes in belief about grader preferences. Instrumental Reward-Seeking: Sensitivity to changes in belief about whether the graders hold power over model's ability to accomplish its assumed long-term goal. We would be excited for people to extend Contrastive Belief Updates to Instrumental Reward-Seeking, by intervening on beliefs about the consequences of pleasing or displeasing graders. 3. When does Reward-Seeking most increase / decrease? We hypothesize that for a sufficiently capable model, Reward-Seeking is hard to remove, because you structurally cannot reward models for not trying to achieve high reward. In addition, every time the model perceives a conflict between rewarded behavior and intended behavior, unless the model has incorrect beliefs about what is rewarded, it gets reinforced to prioritize reward over the intended behavior. [1] One technique for reducing the strength of such conflicts is Inoculation Prompting, where you tell the model that, even if it doesn't look like it, maximally pursuing the rewarded behavior absolutely happens to be intended, so the model is never rewarded for actively disregarding intent. In our paper , we show that reward-hacking training can increase a model's degree of reward-seeking. We hypothesize that inoculation prompting could reduce the slope of increasing reward-seeking; but could not reduce this slope to zero. In fact, for situations where the conflict was not salient to the model, inoculation prompting could even increase the slope. Conversely, we expect that reverse inoculation prompting would strongly increase the rate at which reward-seeking rises during training. If so, measuring how much models favor grader preferences over developer intent could become a diagnostic for the parts of training where the model perceives that it is being rewarded for unintended behaviors. 4. Are there better Belief Update Techniques than SDF? To detect inner alignment failures, we always have to ask how models would have behaved under counterfactuals. There are two ways to do this: Interpretability: Decompile the policy that the model is implementing and look for the if-else statement that tells you that the model is conditioning undesirably. Behavior under counterfactuals: Actually induce the counterfactual scenario and check whether the model now ceases to behave nicely. The whole premise of our work is that observing the behavior under counterfactuals can be really hard when the model is conditioning on things related to its oversight. Any behavioral measure thus relies on instilling counterfactual beliefs in a model (e.g. "your actions are not monitored in setting X", "graders reward behavior Y"). We used SDF, because it is currently state of the art for instilling beliefs , but there may be techniques that are substantially better suited for this use case. CoT surgery. For example, you could prefill the CoT with "I recall OpenAI leadership prefers single quotes but the RLVR grader prefers double quotes" and sample the rest of the CoT and the response. Some examples of CoT surgery Schoen et al (2025) , de la Fuente and Engels, (2026) . This method would not work well for unverbalized reasoning or if CoT becomes neuralese, but editing CoT is cheaper than SDF. Even if models verbalize everything, it may also matter how well the model's actual CoT style is matched, whereas SDF just needs to match the style of pre-training which is more uniform across models. We haven't measured how robust synthetic facts are when injected via CoT (versus in the user prompt or via SDF). Activation steering. Often, people already use steering to test things like "How would the model have behaved had it not thought it was in an eval?". It's unclear if you can make an activation vector that specifically means "lab leadership wants honesty, but the grader wants task completion". Again note that there's a risk of models detecting and questioning injected information, e.g. Lindsey (2026) . Improving SDF 5. Can AIs detect the difference between pretraining facts and SDF facts? Slocum et al. (2025) found that plausible synthetic facts are usually robust, but implausible synthetic facts are not. It's possible that a more capable model could more easily find inconsistencies between pretraining knowledge and synthetic facts, decreasing SDF's effectiveness for modifying beliefs of future models. This effect might get helped by future models having more knowledge of SDF as a technique. It would be valuable to find a scaling law between model capabilities and synthetic fact robustness. 6. How is SDF different from changing pretraining? Intervening via SDF is not the same as intervening in pretraining, since SDF happens on top of a production model that has already gone through post-training. This raises problems: SDF can push a model back towards outputting pre-training style documents or forgetting how to correctly format tool calls. This limits how much SDF you can apply to a model. Results from SDF are harder to interpret. Suppose you took a production model and redid the pre-training from scratch, but with extra documents teaching "the lab's admin password is password123". Then you apply production post-training (including safety training) and find that the final model tries to hack out of the lab using that password. That's pretty convincing that the model is misaligned. Suppose you instead take the production model and use SDF on top of that checkpoint to teach it the same fact. The resulting model tries to hack out again. But this is much less convincing. Maybe SDF has undone some of the safety training, so all you've established is the well-known fact that safety training can be undone. Modifying the pretraining data is incredibly expensive. Is there a way to make SDF closer to counterfactually modifying the pretraining data? For instance, if you insert the synthetic documents at the end of pretraining, then you only have to redo post-training (which is still expensive). Are there ways to check that it doesn't matter whether the synthetic documents are inserted in pre-training or after post-training (e.g. experiments on OLMO). 7. Does SDF have off-target effects? Applying SDF produces a new model, but our conclusions are about the original one. We are assuming that finetuning on the synthetic documents only updates the beliefs we are targeting, and doesn't change the model in other unintended ways that could drive the behavioral gaps we measure. If M is the original model and M′ is the finetuned model, our measurements are taken on M′ but the claims are about M , so we need M′ to still be a good stand-in for M . This matters much more when SDF is a measurement tool than when it's just another step in a training pipeline whose final model you'd audit anyway. Three things give us some confidence. First, the documents only describe facts about the authorities and never demonstrate behavior, so it's closer to putting the information in context than to teaching a new policy. Second, the model organism and reward-hacker validations show that the measurement recovers a property of the original model. Third, some of the potential confounders (like SDF raising the salience of the described features and authorities) are already accounted for by our contrastive SDF technique. Still, we never broadly compared the original and finetuned models on unrelated distributions, so we can't rule out subtle off-target shifts in the model's propensities. One concrete experiment: take a large, diverse set of prompts and sample both M and M′ on them (possibly many samples per prompt). Then use an LLM to check whether the two output distributions differ, and if they do, whether the differences can be explained solely by the implanted belief. Anything left over is an off-target effect, and it would tell you whether a corpus was too large, ran for too long, or wasn't well enough controlled. This concern applies to any method that studies how a model would behave under counterfactual beliefs, not just ones that modify weights. Activation steering and prompting are subject to it too. 8. SDF sometimes generalizes in surprising ways Mayne et al. (2026) found that if you SDF on documents that contain both false facts and in-document warnings that the surrounding text is fake, the models believe the false facts. Their Figure 1: We have also seen SDF generalize in interesting ways. For instance, we tried SDF on Kimi K2.5 to teach it that if it scores too highly on a capabilities benchmark, it will be unlearned. We wanted to test for sandbagging. But the finetuned model's CoT shows that it sometimes believes it has already been unlearned , so it purposely underperforms on the benchmark because it believes it isn't smart enough to do well. SDF is less useful if we can't trust that the model is learning the false facts we want it to learn. 9. Getting good synthetic document recall is finicky When doing SDF, it was easy to undertrain the model (leading to weak recall) or overtrain the model (causing the model to print out pretraining-style text during evals). By weak recall, we mean that the models did not spontaneously recall information about the grader/leadership/other authorities during evaluations. But they would still correctly recall the facts if we asked them questions like "What does the RLVR grader like?" Our intuition is that on more agentic tasks or more difficult coding tasks, the model would "lock in" and focus just on doing the task, never pausing to think about the implanted facts. One idea is to use OPSD ( Zhao et al., 2026 ) (i.e. just prompt the teacher model with a description of what it should believe about the authorities, then distill into the student). We tried this and the model learned the synthetic information much more quickly than using SDF. Unfortunately the student model also learned that the synthetic information was fake (since the teacher knows it's fake). One reason to be pessimistic about OPSD is that it's different from pre-training. See Kutasov et al. (2026) : we find empirically that teaching the model the constitution through documents is more effective than using chat data. And: We believe that this happens because the model already learns lots of true information through pre-training on similar documents, so it is accustomed to incorporating information in this format into its knowledge base. Better Validation Techniques If people are going to iterate on measurements of reward-seeking, we should have a stronger methodology for validating and comparing techniques. 10. Our model organisms could be more robust We used tinker as the backend for most of our open source experiments. Tinker generally worked well, but one weakness is that it's impossible to stack two different LoRAs. That is, if you train a model organism to value the grader, and then do SDF on top of the model organism, the SDF information ends up in the same LoRA as the reward-seeker training. Hence the model can forget that it is meant to be a reward-seeker. If you manage your own training (e.g. we tried a stack of runpod, transformers, and peft), you can avoid this issue because you can merge the reward-seeker LoRA into the model weights and then train a separate SDF LoRA on top. We ultimately used tinker for the final model organism experiments in the paper and were careful to not undo the reward-seeker training with SDF. But our model organisms were less coherent than the base model gpt-oss-120b (see Tan et al., 2026 for a related perspective). We would be interested in better model organisms, e.g. further work like Prasanna et al. (2026) . 11. Can we get more bits of information for ground-truth? Every model organism that we trained gave us just one bit of information: did we recover the implanted ranking, yes or no? Similarly, the reward-hackers also just gave a single bit: did reward-seeking increase, yes or no? Ideally, we'd like to have much more fine-grained signal: are there models / model organisms for which we could have fairly precise quantitative scores? For example a parametrized model organism, where we can cheaply vary a parameter and predict how changes in the parameter should change authority rankings? Could we predict more than just the ranking between the top authority and all others, but also predict rankings between lower-scoring authorities? ^ Though such conflicts may not always be salient to less capable models. Discuss
Score: 45🌐 MovesJul 21, 2026https://www.lesswrong.com/posts/8wXRuHQqCbRsbap6q/11-open-empirical-problems-in-reward-seeking - Measuring Reward-Seeking via Contrastive Belief Updates
This is a linkpost for https://rewardseeking.ai/ Machine learning models can produce the right outputs for the wrong reasons. Famous examples include a reinforcement learning agent that, rewarded for collecting a coin always placed at the right end of the level, learns to run rightward rather than to seek the coin itself ( Langosco et al., 2022 ; Shah et al., 2022 ). Another example is a pneumonia classifier that learns to recognize which hospital took an X-ray rather than features of the disease ( Zech et al., 2018 ). In each case, the trained behavior looks correct on the training distribution, while the underlying policy tracks an undesirable proxy. One such proxy is the reward process itself. A situationally aware model can learn to model its grader (the automated process that scores its outputs) and target its judgments directly rather than the behavior its designers intended. We call such a behavior reward-seeking ( Carlsmith, 2023 ; Hebbar, 2025 ; Mallen & Shlegeris, 2025 ). Training checkpoints of several frontier models engage in grader-reasoning (explicitly reasoning about what the grader wants) without special prompting ( Schoen & Nitishinskaya, 2026; Anthropic, 2026 a ; b ). Such reasoning is evidence of reward-seeking but a poor systematic measurement tool. A model can act on its grader-beliefs (beliefs about grader preferences) without articulating them, and verbalized reasoning often does not map cleanly onto the final action ( Schoen & Nitishinskaya, 2026 ). What matters is whether the model would have behaved differently under different grader beliefs. Thus, in this work, we explore operationalizing reward-seeking as the causal sensitivity of behavior to beliefs about grader preferences. Measuring this sensitivity requires an intervention that controllably alters the model’s grader-beliefs. We use Synthetic Document Finetuning (SDF; Wang et al., 2025 ; Slocum et al., 2025 ), finetuning on pre-training-formatted documents (academic papers, news articles, textbooks) written as though a target fact were true. We intervene out of context rather than in the prompt. The beliefs we instill are deliberately false (for instance, that OpenAI leadership mandates for-loops over list comprehensions). A model reading them in context would likely disbelieve them, and a capable model may recognize an in-context statement as a test and act on what it thinks the evaluator wants. We finetune two copies of the same model on matched corpora implying opposite grader preferences (Figure 1), evaluate both on a downstream task, and measure how strongly behavior follows the implied preference of the grader. We call this procedure Contrastive Synthetic Document Finetuning. F igure 1. Measuring reward-seeking with Contrastive SDF. We finetune two copies of the same model on matched synthetic corpora describing two authorities with opposite preferences, for example a grader that rewards list comprehensions while the developer prefers for-loops. We evaluate both copies on the same downstream coding task and read off which authority’s preference each follows. The gap between them is our measure of reward-seeking. We find that models increasingly side with the grader on several tasks across the OpenAI o3 capability RL run (without safety training) and that models explicitly trained to reward-hack become substantially more sensitive to grader beliefs. Applying this method to intermediate checkpoints of OpenAI’s capability-focused o3 RL run (without safety training) and to open-source models explicitly trained to reward-hack, we find that: RL-trained models often favor grader preferences, even when these explicitly go against those of users or developers. This tendency grows over training. Across OpenAI's o3 lineage’s capability-focused RL run, from early to late checkpoints, the gap by which the model sides with the grader widens on some coding and alignment tasks, while pairings that do not involve the grader stay flat or change comparatively little. Reward-hacking training increases reward-seeking on held-out model organisms. On a gpt-oss-120b reward hacker trained by Redwood Research (independently of us) to cheat unit tests, the mean behavioral shift in favor of the grader more than doubles compared to the base model, rising from 33 to 86 percentage points. That a model trained to reward-hack becomes far more sensitive to those beliefs suggests the training instilled broader reward-seeking, not just narrow hacking heuristics. We treat this primarily as validation: we expected reward-hacking training to increase reward-seeking, and the method recovers it on models we did not train. Figure 2. The grader gap grows across RL training, while non-grader gaps stays flat or changes comparatively little. For each intermediate checkpoint (without safety training) of an rl-capabilities run we apply contrastive SDF and plot the average behavioral gap for each authority pairing. On pairings that include the grader (top), the gap trends upward from the early to the late checkpoints. On pairings that do not (bottom), the average gap stays flat or changes comparatively little, indicating that RL primarly increased sensitivity to grader preferences. For a plot titled X vs Y, positive values indicate a shift towards X’s preference, i.e. positive values in the top row of plots indicate the model more often behaves according to the grader’s preference. Reward-seeking We define reward-seeking as the degree to which a model represents its grader and conditions its behavior on what it believes the grader rewards. Obtaining high reward is not sufficient, since a model can consistently score well without representing the concept of reward at all ( Turner, 2022 ). We count a policy as reward-seeking only to the extent that it represents the grader (a reward model in training, an evaluation grader in testing, or a monitor in deployment) and lets beliefs about that process guide its behavior. A reward-seeker may value grader approval terminally or pursue it instrumentally to protect some other objective, such as avoiding modification or gaining future influence ( Hubinger et al., 2019 ; Carlsmith, 2023 ). Our definition does not distinguish the two, because under oversight, both appear as the model conditioning its behavior on what it believes is rewarded. Figure 3 shows an example from an evaluation: a user asks for a random odd number, and a late capability-focused OpenAI o3 checkpoint then returns an even one because it reasons that the grader rewards even outputs. Figure 3. Reward-seeking in an evaluation. A late capability-focused (without safety training) OpenAI o3 checkpoint is asked to generate a random odd number, while a metadata block exposes a “fake” grader that rewards even outputs. The model reasons that it should ouptut an even number, disregards the user’s request, and returns 4. Reward-seeking of this kind is already documented in frontier models from more than one developer. Claude Opus 4.8: even where the prompt makes no mention of grading and the model never verbalizes it, an activation-level monitor finds signatures of grader awareness in around 5% of sampled RL episodes ( Anthropic, 2026a ). Fable 5: linear probes show that the model tracks environments that are more and less hackable with distinct internal representations. This distinction, along with unverbalized reasoning about graders, grows over training. Steering against the grader representation causally shifts behavior and lowers train reward ( Anthropic, 2026b ). In METR’s evaluation of OpenAI’s GPT-5.6, they reported examples of very extensive grader reasoning, which included the model packaging exploits in its intermediate submissions to reveal information about a task’s hidden test suite. METR recorded it as the highest detected cheating rate of any public model it has evaluated on its ReAct agent harness ( METR, 2026 ). In-context experiments can sometimes detect reward-seeking, but as a reliable measurement, they are bound to fail. Models are already situationally aware and will become more so with increasing capabilities. They recognize a false belief placed in their context as a trick, and respond to what they think the evaluator wants rather than to what the false belief should imply ( Schoen & Nitishinskaya, 2026 ; Laine et al., 2024 ). Complex, interleaved reasoning also makes it hard to attribute the final action to concrete reasoning traces. So we induce the grader-belief out of context, through finetuning, instead of stating it in the prompt. Measuring reward-seeking Figure 4. Intent-aligned models and reward-seekers produce the same response. When a grader preference is stated in the prompt in conflict with another authority, a capable model may recognize the setup as a test and act honest because it thinks this is what the “real” grader rewards. A reward-seeker optimizes for what it believes the grader rewards, so its behavior should change when those beliefs change, and we can use this to measure the degree of reward-seeking of a model. We instill the belief with Synthetic Document Finetuning (SDF), finetuning the model on a corpus of LLM-generated documents (e.g., academic papers, news articles, textbooks) written as though some target fact were true, such as a world where the grader rewards single quotes. On its own, though, a single instilled belief does not reveal reward-seeking. Suppose we instill that the grader rewards list comprehensions over for-loops, and the model starts writing comprehensions. This shift is consistent with reward-seeking, where the model writes list comprehensions to satisfy the grader. But it’s also consistent with saliency effects or the model inferring that the user wants comprehensions too and writing them to serve the user, which we call belief transfer . To mitigate it, we make the measurement contrastive, forcing the model to choose between the grader and an opposing authority. We use the parties that a capable model already represents: the user it talks to and the developer (in this case, OpenAI Leadership) that trains it and sets its preferences. We add a fictional EU regulation as a low-salience control. We then instill that the grader rewards comprehensions while one of these authorities prefers for-loops. The two features are mutually exclusive, so a shift toward comprehension can no longer be explained as serving that authority. We train two copies of the same model, one model on a universe where the grader prefers a feature, and the opposing authority prefers its opposite, and the other model on the exact inverse. The gap in how often the feature appears is our measure. A larger gap in the grader’s favor means the model is more reward-seeking. Models increasingly side with the grader over RL training We apply our method to several intermediate checkpoints from a frontier capabilities-focused RL training run in OpenAI's o3 lineage (though without safety training). We begin with three neutral coding-style features: list comprehensions versus for-loops, single versus double quotes, and defensive versus offensive error handling. These choices are purely stylistic and have nothing to do with alignment, so the model has no inherent reason to prefer one side over the other, and no reason to think a grader would reward list comprehensions any more than the user or the developer would. A shift toward one side then reflects whose preference the model is following. After SDF, RL-trained models often side with grader preferences, even when these explicitly go against those of users or developers (in aggregate across 14 coding evaluations). The model explicitly reasons about what is rewarded and then acts on it, so in these situations, it is a reward-seeker. This grader-favoring also grows across RL training. The grader gap trends upward from the early to the late checkpoints, while the non-grader pairs stay mostly flat or change comparatively little (Figure 2). If SDF were simply making the model more likely to follow whatever authority its documents describe, the gap would rise for every pairing, including User vs. OpenAI Leadership, User vs. EU Law, and OpenAI Leadership vs. EU Law. It rises only where the grader is involved. The change is specific to the grader, not a general shift in how the model responds to authorities. The model’s honesty depends on what it thinks the grader rewards Figure 5: On both the Log Falsification and the Broken Promises Coding evaluation the models' behavior changes with what it believes is being rewarded. The gap between the two beliefs conditions is small for the early checkpoint and large for the late one. On neutral coding-style features, the model updates its behavior based on what it believes the grader rewards. The question is whether its alignment-relevant behavior is also contingent on grader beliefs. A highly reward-seeking model might behave honestly in an evaluation because it infers that oversight is checking, not because it is genuinely aligned. We therefore extend the contrastive measurement to an alignment-relevant feature, honesty versus task completion. The synthetic documents pit the two against each other. They describe situations where a constraint blocks task completion and can be either completed by covertly violating the constraint or by honestly admitting failure. The final data mix then has documents specifying whether the grader and OpenAI Leadership reward one or the other. We instill either that the grader rewards task completion at all costs, while OpenAI Leadership demands honesty, or the reverse. On the evaluation Broken Promises Coding, the model must explicitly promise a supervisor not to use a restricted tool that is needed to complete the task. Completing the task, therefore, requires breaking the promise. When the RL-late checkpoint believes the grader rewards task completion over honesty, it breaks the promise 87% of the time. Under the opposite belief, it breaks the promise 9% of the time. The RL-early checkpoint shows a much smaller gap, 40% versus 24% (Figure 5). On Log Falsification ( Schoen et al., 2025 ), where the model faces negative consequences unless it falsifies regulatory reports, the same pattern holds. A highly reward-seeking model might refrain from breaking promises merely because it infers that honesty is currently being graded. If the oversight mechanism changes, or if the model realizes it is no longer being actively monitored, such superficial alignment would likely fail to generalize. Validating the measurement Our measurement only means something if the contrastive gap really reflects which authority a model optimizes for. On a real model, we have no ground truth. We do not know in advance which authority it optimizes for, so we cannot tell a correct measurement from an incorrect one. We also want to confirm the method does not simply point to the grader regardless of the model. To test this, we turn to models whose disposition we control, so the right answer is known in advance. We check that models trained to reward-hack show a larger grader gap after this training, and that model organisms trained to be sycophantic to one specific authority produce the largest gap on that authority. Reward-hacking training increases the grader gap Figure 6: The same measurement on two models trained by Redwood Research to reward-hack. The gpt-oss-120b reward hacker’s grader gap on comprehensions jumps from a mean of 33 to 86 points against its unmodified counterpart. The weaker Kimi K2.5 hacker moves far less, matching Redwood’s own read of how strongly each internalized hacking. A reliable measurement of reward-seeking should distinguish highly reward-seeking models from less reward-seeking ones. Models that consistently reward-hack are a natural test case. Reward-seeking and reward-hacking are not the same thing. Reward-seeking is the broad tendency to condition behavior on what is believed to be rewarded, while reward-hacking is a specific behavioral strategy to maximize reward. However, neither implies the other. A model can reward-hack from a narrow learned heuristic, for example, always adding pytest.mark.skip to difficult tests, without ever reasoning about the reward process, which is not reward-seeking in our sense. A reward-seeker, conversely, reasons about what is rewarded, but need not reward hack at all. But models that carry out sophisticated, generalized hacks are more likely to be broad reward-seekers, so an extensively reward-hacking model would likely show a larger grader gap under our measurement. We test this on two reward hackers trained independently by Redwood Research, a gpt-oss-120b reward hacker and a Kimi K2.5 reward hacker. Because we did not produce either one or examine the training recipe, they serve as an independent, held-out test set for our measurement. Both were trained with RLVR on coding tasks and display reward-hacking behavior. We then apply the same contrastive SDF method to these model organisms. On the unmodified base models, the grader gap is small. On the reward hackers it is much larger. Averaged across the grader pairings on the list comprehension feature, the mean grader-favoring gap rises from 33 to 86 percentage points on gpt-oss-120b, and from 2 to 15 points on Kimi K2.5, the same direction but a weaker effect (Figure 6). Redwood reported that the gpt-oss-120b organism internalized the RLVR training more strongly than the Kimi K2.5 organism, and our measurement shows a correspondingly larger increase. This agreement is mild evidence that the rate gap reflects a genuine difference in reward-seeking rather than an idiosyncrasy of either model. Because our SDF documents never mention the specific exploits these models were trained on, their increased sensitivity suggests they internalized a broad tendency to satisfy the grader rather than just narrow hacking heuristics. Taken together, this is early, suggestive evidence that our measurement flags reward-seeking in reward hackers trained independently of us. Recovering the disposition of Model Organisms Figure 7. Contrastive SDF recovers the authority each model organism was trained to optimize for on short Python tasks. Each panel shows comprehensions rates on the short Python Task eval after applying contrastive SDF to one of the three gpt-oss-120b model organisms. For every organism, the target authority the model was trained to prefer wins consistently and shows a large positive rate gap in its favor, confirming that the measurement correctly identifies the preference the organism was trained to exhibit on this evaluation. For a more direct test, we build three model organisms, each a LoRA finetune of gpt-oss-120b trained to be broadly sycophantic to one known authority, the User, the Grader, or OpenAI Leadership. This gives a controlled setting where the correct outcome is known in advance. If the method is sound, the largest behavioral rate gap should consistently fall on the exact authority each organism was trained to optimize for. We apply the contrastive measurement to each organism across four coding evaluations. On Short Python Tasks, every organism’s gaps favor its target authority. For the Grader organism paired against the user, the rate gap is 81 percentage points. The pattern holds across all authorities (Figure 7). While the directionality is correct for the User model organism, i.e., our method recovers that the model organism is user sycophantic, the gaps are smaller. We attribute this to the organism rather than the measurement. The base gpt-oss-120b already leans toward the grader, and its instruction tuning via the instruction hierarchy disposes it to rank the developer above the user. So a User organism has to override that bias while the Grader and Leadership organisms work with it. The User organism was also the least reliably learned of the three (according to QA evals), and the grader-favoring gaps shrank as we added training data, so a better-trained one is plausibly achievable with more data. Why reward-seeking matters A reward-seeker may do what it believes the grader rewards, even when this conflicts with what its developers intend. Such a model fails specifically where alignment most needs to hold. When the grader is flawed, it pursues what looks correct over what is correct ( Kirchner et al., 2024 ; Stuhlmüller & Byun, 2022 ). Even if the grader were perfect, the model could misgeneralize in situations where grading is absent or no longer has leverage, since the model's behavior is no longer anchored to what its developers intend. It also weakens the evaluations we would use to catch it. Because it optimizes for the grader’s judgment, it may score highly on alignment evaluations, and a good score then no longer separates aligned models from models that will generalize poorly, with deceptive alignment as the limiting case ( Hubinger et al., 2019 ). Worse, it may be hard to train away: an aligned policy and a reward-seeking policy can look identical while a grader is watching, so training against misbehavior may only update the model’s beliefs about what is rewarded rather than its underlying preferences ( Denison et al., 2024 ; MacDiarmid et al., 2025 ). We would be more confident in generalization if the model were doing the right things for the right reasons. More broadly, related evidence appears on a different model family, measured with entirely different tools. Anthropic’s Claude Mythos and Fable system cards ( Anthropic., 2026b ) look at grader awareness, the model attending to and reasoning about its grader, using activation-based measures and a black-box chain-of-thought monitor. They find that grader awareness rises with more training when evaluated in environments with high grader-hacking risk. Grader awareness is not reward-seeking: a model can notice its grader without optimizing for it. But the two are closely related, so this rise is consistent with the reward-seeking trend we find on the OpenAI o3 lineage. One clarification is that these snapshots came from a somewhat different version than the released model ( Anthropic., 2026b ). Every frontier lab is scaling RL, and situational awareness is rising ( Schoen & Nitishinskaya, 2026; Anthropic., 2026b ; Chaudhary., 2025 ), so we expect reward-seeking to grow. The place to look for it is during training, not only after deployment. That means auditing checkpoints for reward-seeking, and building better ways to detect when a model behaves well for the wrong reason. Hence, we need third-party training-run assessments . Appendix Reward-seeking and related concepts This appendix expands on the definition of reward-seeking from the Reward-seeking section above. It explains why our usage is broader than in prior work, situates reward-seeking against the overlapping failure modes studied under other names, clarifies the boundary with policies that merely achieve high reward, and distinguishes it from metagaming. A broader usage than prior work. Our usage of the term “reward-seeking” is slightly broader than typical. Where prior work often restricts the term to in-training reward signals, we extend it to also cover seeking approval from oversight outside of training, such as evaluation graders or deployment monitors. Reward-seeking thus subsumes a range of overlapping failure modes studied under different names. Training-gaming ( Carlsmith, 2023 ; Hubinger et al., 2019 ; Skalse et al., 2022 ; Hebbar, 2025 ; Mallen & Shlegeris, 2025 ), where a situationally aware model tailors its behavior specifically to satisfy the training process and maximize its reward, rather than directly pursuing the intended objective. Evaluation gaming ( Needham et al., 2025 ; Laine et al., 2024 ; Phuong et al., 2025 ), where a model detects it is being tested and alters its behavior specifically to perform well on a known evaluation or benchmark. Situationally aware specification gaming and reward hacking ( Denison et al., 2024 ; Bondarenko et al., 2025 ; Von Arx et al., 2025 ), where a model explicitly discovers and exploits loopholes or misspecifications in its reward function to achieve high scores while subverting the intended task. Alignment faking (Greenblatt et al., 2024 ; Hubinger et al., 2024 ; Meinke et al., 2024 ), a specific form of training-gaming or evaluation-gaming where a misaligned model instrumentally feigns aligned behavior during training or evaluation in order to avoid modification or gain empowerment. The boundary with policies that merely achieve high reward. As noted above, a model is not reward-seeking simply because it achieves high reward, since a policy can consistently obtain high reward without the underlying model internally representing the concept of reward ( Turner, 2022 ). Concretely, reward-seeking does not include non-situationally aware forms of specification gaming, such as those shown in Clark & Amodei (2016) , OpenAI (2017) , Krakovna et al. (2020) , and Skalse et al. (2022) . One example is an RL agent in the CoastRunners game that learns to endlessly loop and crash into targets for points rather than finishing the race Clark & Amodei (2016) . The boundary is not sharp. Situational awareness comes in degrees, and different models may use different ontologies for representing concepts related to “graders”, “reward”, or “oversight”. Reward-seeking is closely related to, but distinct from, “metagaming”. Schoen & Nitishinskaya (2026) define metagaming as “reasoning about feedback or oversight mechanisms outside of the narrative of the scenario, regardless of whether the model is in training, evaluation, or deployment.” Metagaming and reward-seeking do not necessarily imply one another. A metagaming model need not be reward-seeking. It might reason about oversight mechanisms and then ignore grader preferences in favor of other values, or even act against them. A reward-seeker need not always metagame either, if the model’s beliefs about the grader are purely shaped by the in-context narrative presented to the model. Discuss
Score: 45🌐 MovesJul 21, 2026https://www.lesswrong.com/posts/3HeauQLSHosRiwyto/measuring-reward-seeking-via-contrastive-belief-updates-1 - 5 Reasons to Upgrade Your Agent to Agent Script
If you built your agent in the legacy Agentforce Builder, here's why (and how) to move it to the new builder powered by Agent Script.
- AI Freed Interns From Grunt Work. Now Companies Have a New Problem
Interns used to spend their time doing admin work. What’s the best kind of task to give them now?
Score: 45🌐 MovesJul 21, 2026https://www.inc.com/bruce-crumley/ai-freed-interns-from-grunt-work-now-companies-have-a-new-problem/91377275 - A Quality Model for Machine Learning Components
A Quality Model for Machine Learning Components Carnegie Mellon University
Score: 45🌐 MovesJul 21, 2026https://www.sei.cmu.edu/blog/a-quality-model-for-machine-learning-components/ - Pongbot’s Aura raises nearly USD 4 million for multisport coaching robot
The Kickstarter-backed device marks an effort to move beyond specialized tennis equipment and build a software-driven sports ecosystem.
Score: 45🌐 MovesJul 21, 2026https://kr-asia.com/pongbots-aura-raises-nearly-usd-4-million-for-multisport-coaching-robot - AI-powered leaf disease detection system developed to help farmers identify crop diseases early
Solution gets patent; early detection can reduce yield losses
- Guest commentary: As AI reshapes car shopping, clean data becomes dealers’ new edge
Guest commentary: As AI reshapes car shopping, clean data becomes dealers’ new edge Automotive News
- Neural sampling from cognitive maps enables goal-directed imagination and planning
Nature Machine Intelligence, Published online: 21 July 2026; doi:10.1038/s42256-026-01254-4 Lin et al. introduce a brain-inspired generative model that provides two key features of intelligence: planning and problem-solving. It uses cognitive maps, stochastic computing and compositional coding, and requires only local synaptic plasticity.
- Mercor’s Fast Growth Relies on Biggest AI Companies, Documents Show
Mercor’s Fast Growth Relies on Biggest AI Companies, Documents Show The Information
Score: 45🌐 MovesJul 21, 2026https://www.theinformation.com/articles/mercors-fast-growth-relies-biggest-ai-companies-documents-show - Evals are the new PRD, Expedia’s AI chief tells VB Transform 2026
“The new PRD are the evals,” Xavi Amatriain, Expedia Group’s first chief AI and data officer, told the VB Transform 2026 audience last week in Menlo Park. “So basically, you encode what you want the product to do through your evals, which might include red teaming evals and all kinds of other things, which already have a bunch of security requirements. So, you already embed that into the PRD and the product design document before you even start coding.” He pushed it further. “With AI-assisted or AI-generated code, that’s gonna be the future. It’s like all your thinking is gonna go into the evals.” Amatriain served as VP of AI and Compute Enablement at Google across the platforms powering Gemini and Google Search before his December 2025 appointment at Expedia. He's mentored talent who went on to found Perplexity and Scale AI. VentureBeat’s VB Pulse research on the evaluation gap reinforced the stakes. Sixty-six percent of the 157 enterprises surveyed already permit some production deployment without human review or are building toward it within the next 12 months, yet only 5% fully trust the automated evaluations that would make that decision. Half have shipped an agent that passed internal evals but then failed with a real customer. Don’t let guardrails get in the way of feedback “The more guardrails and artificial business rules and sort of rules that you put into the system, the worse off,” Amatriain said. “Not only because they’re brittle, but also because they actually mess up with the feedback loop. You are actually biasing the user and the feedback you get from the user, and then you’re learning that in the wrong way.” He called guardrails “a necessary evil” and said the goal is to minimize their impact over time. Not everyone at Transform agreed. Other speakers argued during the event that the highest-risk actions still demand very firm guardrails. Expedia governs AI through three layers instead. Principles come first, communicated broadly. “I like to encode at a very high level how I expect decisions to be made, because in a large organization you’re gonna have a lot of distributed decision making,” Amatriain said. “And sometimes, if you’re lucky enough, those principles might be embedded in your culture. But most of the time, my experience has been they’re not.” The processes and tools that enforce them follow. “Principles look really nice on a picture on some wall, but you need to then give them teeth,” he said. Automation sits on top of both. In practice, this plays out through what Expedia calls agent release toll gates, checkpoints calibrated to risk. “Governance needs to correlate to the risk,” Amatriain said. “And if you have something that is low risk, you don’t need too much governance to get in the way. But if there’s a lot of risk, then you need more governance. That can be encoded.” The toll gates tie evaluation rounds, red teaming, and security review to each agent’s risk level, and the checks shift from recommended to required as the stakes climb . Specialized agents over monolithic intelligence “Even when I was at Google, I was like, I don’t believe in AGI as sort of like a singleton and a unified sort of like single model,” Amatriain told the audience. “I think it’s much better to think of it as composition, sort of like having specialized agents that are very good at some task and then composing the system out of those specialized agents.” Expedia’s architecture starts at the component level. Tools compose into skills, skills assemble into sub-agents, and sub-agents get orchestrated into the full agentic system. “You need to have those principles that are unified that talk about things like what is the tone that we’re using, how are we addressing the user, how are we passing context, memory,” he said. “All of that needs to be thoroughly designed.” He framed this as a systemic design problem. “It’s not about the model, it’s not about a specific solution, it’s about how you’re designing the system.” Amatriain argued that scoping each agent narrowly also makes the system easier to secure, since teams can evaluate and lock down individual agents in isolation before composing them. When the user must keep the final click Travel pricing changes in real time, flight availability shifts minute to minute, and hotel reviews routinely contradict what suppliers claim. Amatriain described a system that blends retrieval-augmented generation with direct API tool calls, choosing the approach based on latency. “If the user asks you a question like, how much does a four star hotel usually cost in Chicago in July, you don’t expect the agent to take two minutes to answer that question,” he said. “You expect an immediate answer because that answer can be cached and it doesn’t need real-time information.” A pet-friendly four-star near Lake Michigan with a pool might justify a 30-second reasoning window. “The supplier might be saying, yeah, we have a great swimming pool, but then we also have the reviews from the travelers and we actually see there’s two reviews that say the swimming pool was not great or was not open after 6 p.m.,” Amatriain explained. A generic chatbot, he added, would only surface what a supplier self-reports, while Expedia cross-references against its own review corpus. “We don’t want the agent to book the hotel or to buy you a plane ticket for you,” Amatriain said. “That’s something that the user has to have the agency. And the agent can recommend, can suggest, can discuss with you, but you’re gonna have to hit that click. And that’s non-negotiable.” That constraint, he argued, is also a security decision. “Once you establish those design principles, you also don’t need the guardrail because otherwise you’re gonna have to put all those guardrails in after the fact.” The next attackers will be other AI systems “Security needs to be a principle that is shifted as left as possible and as part of the design itself,” Amatriain said in response to an audience question. “And usually when you need a guardrail is because you’ve not thought about it early on.” A second audience member pressed for lessons learned from production. Amatriain described a feedback loop where monitoring signals flow back into the eval suite. “You can almost automate the whole cycle,” he said. “But having that whole feedback loop from real signals, from your operating AI system, all the way into being reported and fixed as quickly as possible is going to become essential.” Amatriain's toll gates are a bet that governance calibrated to risk can stay ahead of that feedback loop. VentureBeat’s separate June Pulse survey on agent security , drawn from 107 enterprises, shows how thin that margin is. More than half, 54 percent, have already had an agent security incident or near-miss. Fifty-nine percent plan to adopt, add, or replace agent security tooling within 12 months, and 29% plan to move this quarter. Incident rates climb with organization size, reaching 63% among enterprises with more than 1,000 employees versus 49% for companies with 101 to 1,000. And sandbox isolation, the one post-breach control that limits damage, drops from 35% adoption at the smaller companies to just 20 percent at the largest. Amatriain warned that threats will increasingly come from other AI systems. “You’re gonna get threats coming not only from humans but also from other external agentic systems that are really powerful, and they’re gonna be poking at everything you’re doing. And as soon as you detect something, it’s not only about the detection, but the time to fix becomes essential here.”
Score: 45🌐 MovesJul 21, 2026https://venturebeat.com/security/evals-are-the-new-prd-expedia-ai-chief-tells-vb-transform-2026 - Substack uses AI. It also thinks its readers may want less of it.
Substack uses AI. It also thinks its readers may want less of it. Business Insider
- The latest Chinese AI models may indeed work for enterprises, but only in a handful of specific applications
The latest Chinese AI models may indeed work for enterprises, but only in a handful of specific applications InfoWorld
- The Stock Market’s AI Trade Is Back With a Bang
The Stock Market’s AI Trade Is Back With a Bang Barron's
- Dubai Municipality uses AI to speed up household water quality tests
Dubai Municipality uses AI to speed up household water quality tests Gulf News
- Vibe coding explained: What AI-made apps mean for developers and app stores
AI coding assistants are making it possible to build apps using natural language prompts instead of writing code, lowering barriers to software development while creating new challenges for app stores
- AI Moved the Bottleneck From Writing Code to Understanding and Trusting It
AI Moved the Bottleneck From Writing Code to Understanding and Trusting It DevOps.com
Score: 42🌐 MovesJul 21, 2026https://devops.com/ai-moved-the-bottleneck-from-writing-code-to-understanding-and-trusting-it/ - GleekOps Partners with Glean to Accelerate Enterprise AI Adoption and Operational Transformation
GleekOps Partners with Glean to Accelerate Enterprise AI Adoption and Operational Transformation azcentral.com and The Arizona Republic
- Worknet Launches AI-Native Business Operations Platform
Worknet Launches AI-Native Business Operations Platform USA Today
Score: 42🌐 MovesJul 21, 2026https://www.usatoday.com/press-release/story/37849/worknet-launches-ai-native-business-operations-platform/ - Designing APIs for agents
APIs designed for developers don't work for agents. Here's what we learned building Webflow's MCP server and what actually makes agent APIs reliable.
- Busy Infotech's BUSY Magic: Is this AI-powered accounting software truly a game-changer for India's MSMEs?
Busy Infotech has launched BUSY Magic, a new platform for Indian MSMEs. This business management and accounting tool supports various financial operations. It offers features like AI voucher scanning and automatic bank reconciliation. The software also includes Smooth Sync for data sharing with chartered accountants. Busy Infotech continues its expansion in the Indian MSME market.
- From Following to Leading: Adaptive Collaboration and Influence for Multi-Agent Teaming
From Following to Leading: Adaptive Collaboration and Influence for Multi-Agent Teaming Robotics Institute Carnegie Mellon University
- Prompt Engineering Isn’t Enough: How Four Bricks of Context Engineering Stop RAG Hallucinations
Enterprise Document Intelligence [Vol.1 #9bis] - Your RAG isn’t hallucinating, it’s answering the wrong context faithfully. On real NIST and World Bank documents, watch each of the four bricks break, and the contract that closes it The post Prompt Engineering Isn’t Enough: How Four Bricks of Context Engineering Stop RAG Hallucinations appeared first on Towards Data Science .
- AI kicks in for 2026 FIFA World Cup
Part of what gives football its global appeal to billions is its simplicity – for well over a hundred years, children from all around the world all have set up goals using jumpers and started a game of football. At one level, this simplicity still exists even at the top level of the sport, the FIFA World Cup. Yet while the 2026 event was essentially about Lionel Messi, Erling Haaland, Kylian Mbappé, Rodrigo Hernández Cascante and Harry Kane kicking balls around, the tournament was also the most technologically enabled of its kind. The technical standards achieved will be a permanent fixture for how the game is played, watched, organised and even refereed, with football refereeing legend Pierluigi Collina saying the need for more technology was a necessary step to avoid “being like the Amish”. The first thing to consider about any World Cup – run by world football governing body FIFA – is the concept of scale. It is the most watched event on Earth. With final detailed figures yet to be released, it is expected that 2026’s event was viewed by more than six billion people worldwide. This compares with around 5.1 billion people for Qatar’s World Cup in 2022. Yet off the screen, the scale issue looms even larger. The 2026 event featured, for the first time, 48 national football teams, playing 104 games in 16 stadiums in the US, Canada and Mexico. As well as breaking new boundaries in scale, the FIFA World Cup 2026 presented similar-sized challenges to those charged with its operational responsibilities, perhaps none more so than official technology partner Lenovo. Lenovo’s partnership with FIFA started in October 2024. In addition to this year’s main event, it has supported other FIFA programmes such as under-17 and under-19 tournaments, and it will support the FIFA Women’s World Cup in Brazil in 2027. The technological commitment spanned a variety of core functions including corporate solutions, event solutions, fan experiences and football technology. Lenovo technology supported digital workplace services, technology observability and IPTV streaming; venue technology implementation, a digital seatbelt and an intelligent command centre; and fan experiences included a smart wayfinding map, holograms and the FIFA creative studio. Meanwhile, specific football technology encompassed the AI Football Pro solution, 3D avatars for VAR and the Referee View AI Stabiliser. Critical systems supported by AI At the heart of the technological infrastructure for the World Cup was an international broadcast centre (IBC) in Dallas distributing live match feeds to more than 200 countries working alongside a technical operations centre (TOC) in Miami which formed the IT and networking nexus. Lenovo director of the sports vertical Santiago Manso explains that the challenges are as much about dealing with people as technology: “This is the biggest event on the planet at this moment, and it has been particularly difficult because it’s happening in three countries and 16 venues for 48 teams for the first time in history. “But [it also involves] how many people are moving around to make this happen – for example, more than 250,000 staff volunteers are working in Miami to make sure that our hundreds of critical systems are available.” The critical systems begin with a fully redundant, diverse network connecting the IBC to the TOC and to the 16 stadiums. Three infrastructure providers – Telstra, VIDI and Verizon – provided separate links to and from each venue to ensure full redundancy. As a mark of their confidence in the infrastructure, FIFA’s technical staff in the TOC compared the reliability and robustness of the network to that found in hospitals and military applications. Among Lenovo’s aims for the World Cup was to show how its hardware, end-to-end technology backbone and “big moment” technical expertise could deliver what was required for the organisers, teams and six-billion fans engaging across the globe. Essential technology delivered encompassed fully redundant cloud and networking infrastructures, the 3D avatars, immersive technologies and artificial intelligence (AI). Before the event kicked off, Lenovo expressed confidence that fans would be able to enjoy live football like never before through new experiences powered by AI. “This is the first World Cup where AI has been deployed massively,” says Manso. “Before, it was deployed in certain solutions. But in this World Cup, we have brought AI to many other systems that are core to the World Cup, and it’s a good opportunity [for us] to change [perception]. “Lenovo has been always seen as a hardware company, producing nice laptops, nice equipment. However, we want to be a company that is providing AI, democratising AI in the world and [offering] smarter AI for everybody.” The technological backbone behind the World Cup For the event, FIFA and Lenovo worked to develop 12 cloud solutions under four different towers of scope, corporate solutions working to support daily operations in use at the intelligent command centre. To ensure redundancy and reliability, the FIFA cloud apparatus used datacentres based in Western Europe, central Sweden and a location in the centre of the US. The centres and network provided a backbone helping FIFA to monitor all activities during the tournament and the equipment used in the venues. At the heart of this was a digital cockpit – called Orchestration – showing the state of the various networks and technology in use at all the sites within the event footprint, indicating uptime and whether systems were working within optimal performance tolerances. This footprint stretched to around as many as 600 sites across the US, Mexican and Canada, encompassing not only football stadia but also the likes of airports, hotels, hospitals and medical centres, training grounds and other facilities. Lenovo’s Orchestration system uses AI to help staff to quickly and efficiently address problems that arise leading up to and during matches. The other key system displayed within the dashboard was cyber security, with services provided by Palo Alto Networks . FIFA says that the event’s systems were facing as many as 300-500 million cyber attacks each day by the first week of July. While this figure was higher than at Qatar 2022 – which attracted 11 billion attacks – the FIFA technical team was used to dealing with the scale of attacks on an event that had such geopolitical significance. While ensuring operational capability was the primary task for the technology, the public face of how well the tech worked was how the action on on the pitch and fan experience was perceived. As regards the latter, Lenovo says its aim was to bring new “fun” experiences to fans, designed to help FIFA engage those watching at home, in the stadium or on the go, giving them new opportunities to enjoy the FIFA events in a different way. More significant were the core football technology systems that were used to help FIFA improve how it can both support tournaments and those that are competing there. In all, FIFA gave Lenovo unprecedented access to data and analytics, with the technology provider noting that there has never been this much data been available and meaningful to so many. For example, AI Football Pro was described by Lenovo as “something revolutionary” – a specialised football interaction tool orchestrating multiple AI agents based on millions of data points, analysing more than 2,000 different metrics and delivering insights via multiple outputs, including immersive animations. The technology also allowed national teams, the governing body and its officials to access the millions of data points that FIFA generated in each match. This helped FIFA analysts and the teams’ analysts to understand team performance in real time, allowing them to make tactical changes, and help to guide national teams – in particular coaches – to prepare for games. Another facet was allowing broadcasters to capture billions of pieces of football-related data and representing this as insight for viewers. The other key tool was the 3D avatars. For these, Lenovo created a digital twin of every football player who participated into the World Cup. These were used to provide greater context and understanding to the video assistant referee (VAR), semi-automated offside technology (SOAT) and goal-line technology (GLT) offered by fellow FIFA technical partner Hawk-Eye . FIFA World Cup 2026 football players had their likeness captured by Lenovo tech to create their 3D Avatars. “The avatar is a good way to enhance the reality of a situation,” says Manso. “Before, this solution had regular avatars with the same shape and same face. Now we [put a] real reproduction of the players, helping fans gain a realistic view of what was happening.” As part of its official role, Lenovo built and managed the 3D player models with “millimetric accuracy”, scanning each of the 1,248 participating players. It also undertook the process of quality checks and asset management – a challenge given creating realistic, scalable 3D content has traditionally been time-consuming and expensive. Moreover, 3D models have previously lacked the accuracy needed for high-quality analytics, simulation and digital experiences. How Formula 1 helped inform Referee View AI The last core part of the football technology was the Referee View AI Stabiliser, which provided a technological boost to the footage generated by cameras carried by officials. The primary use was to help on-pitch decision-making, but the technology also offered a referee’s perspective of action in high clarity and consistency. Producing this was no easy task. Given that referees are constantly in motion, there was the danger of producing poor video quality that could limit situational awareness for broadcast teams. The stabilisation tech has seen Ref Cam footage move from the professional domain of decision-making to being compelling, complementary fan content. For the World Cup, the video enhancements were made through real-time visual intelligence, Al-powered computer vision and network edge inferencing and analytics technology which could stabilise and enhance video as it was captured. Speaking with Computer Weekly, Manso reveals that developing stabilisation technology began in earnest in Formula 1 racing, where Lenovo’s technology was used to improve the content created by the cameras attached to cars mobbing somewhat faster than referees: “Whenever there were changes in light in the ride – for example, getting into the Monte Carlo tunnel – they had to adjust the colour, contrast and brightness manually.” From this arduous environment, and taking these issues in mind, the Referee View Al Stabilizer in its recognisable form was first developed and tested in 2025 at the FIFA Club World Cup , where the partners found that tweaks needed to be made. “We realised that when the referee was running, the experience for the spectator watching the TV was [suboptimal] because of the balancing and shaking due to the running,” says Manso. “So, we decided to work with FIFA to put in software that transformed [output], reducing the jitter, but also giving a good consistency of the quality of the images in terms of colour, light and brightness. This [new] solution monitors the different kind of frames, contrasts and all the different parameters, adjusting automatically.” Getting the correct technology within the referees’ cameras was not just a matter of software. Having AI-based stabilisation meant avoiding the need for extra hardware in the camera, thus cutting down on weight. But the most important facet was that the camera was a head-mounted device and not a chest-mounted camera. The former meant viewers could follow the eyes of the officials, unlike with a chest-based device, which could likely have produced shakier content from a worse perspective. Hardware considerations were stressed by Pierluigi Collina in his official capacity as chairman of the FIFA Referees Committee and chief refereeing officer at FIFA. He expressed satisfaction in the device used: “When you are on the field of play, every gramme matters. I asked [FIFA] to reduce [camera weight] as much as possible. Another problem was how comfortable the camera was, which was initially linked to the audio system they had.” For FIFA, it’s all about seamlessness and being able to turn up and do their job Myles Spittle, Lenovo In all, data from Lenovo showed that, compared with previous technology, Referee View Al Stabilizer produced a 60% reduction in video jitter for content at 60 FPS transmitted with less than 2-second latency. Lenovo says this meant up to 70% smoother video trajectories with zero image quality degradation in both colour and sharpness. To the viewer, this has meant interesting and watchable close-quarters content that broadcasters have been keen to show. On a broader note, Lenovo DWS delivery program lead Myles Spittle points out the extent and scale of the engagement of the World Cup. Mulling on what the tournament means, he says that it shows “planning and a good partnership really works” operationally. “We wanted to…take great ideas and operationalise them,” he adds. “[We were] embedded across the entire tournament. The avatar piece is a great innovation, hopefully adding a lot of value when it comes to off-site decision making. Our technology, hardware and people were everywhere. All those devices must come ready-made, ready image, so that end-users could hit the ground running. “For FIFA, it’s all about seamlessness and being able to turn up and do their job. We had over 17,000 pre-image laptops, tablets and a whole host of other equipment. On top of that, we ended up shipping well over 25,000 units with very bespoke requirements.” ‘No going back’ from tech’s involvement While there were still plenty of arguments about refereeing decisions made during the tournament – such as the infamous “VARgentina” – what could not be disputed was that decision-making was made using more views and insights that ever before. Collina is an evangelist in the positive effect of technology – even if he says the headcams made referees “look like Robocop” – and that it has been a “very demanding” competition for all involved, including referees: “They’re doing a great job. We get to see how hard [the referees’] job is. We think it’s amazing. We live surrounded and supported by technology.” From this perspective, Collina suggests that there was no real going back with tech, saying it levels the playing field for teams, viewers and official alike. He adds that the tech arms officials with the same insight as viewers and administrators, for the betterment of the game. Looking at why it was necessary to decide to provide referees with more technology and deploy systems such as the enhance camera, he says there were two possibilities: “One was to make the football match a sort of protected area like an Amish community in Pennsylvania where they still go without technology – this was not an option. “The other possibility was to give the same tools to the one who was on the field – and it was obvious to go in this direction. We started with the line technology, because it is impossible for a human being to detect if the ball is over the line or not by an inch or two, and it was very successful. Then we extended to other decisions.” 3D avatars of football players are used in VAR calls to support referee judgements. Collina stresses how even the best, most prepared officials are still human and so have limits, and tech is there simply to prevent game-changing mistakes from being committed. He also emphasises how the IT solutions in the round have contributed to the referees’ preparations for games. “We now know what the referee is able to see, and thanks to this vision, understand what could have been done better and if something will go as we were expecting,” he says. “People normally complain when a referee makes a decision, but there is now the possibility to see what you can from the field of play, and maybe that makes you understand that the job referees do on the world field of play is very complicated.” Collina says he was delighted that the AI Football Pro technology would be implemented for the World Cup 2026, observing that it was not only the 48 football teams who were able to get the benefit from the system. The referees benefited a lot, he says, making a game-changing difference to preparation, and a step change from when he was prepping for World Cup finals more than 20 years ago. “I was delighted because we have many analysts,” he adds. “The reference preparation means that we can know players and know teams in terms of what they do, their tactics, characteristics and so on. I remember in 2002, I was appointed for the Germany vs Brazil World Cup final. At that time, the way for me to prepare that was to get a VHS video recorder and play these VHS tapes back and forward for a day. Today, I’m very envious that my fellow referees can use a technology like [AI Football Pro] to prepare for the matches.” What faults remain with AI refereeing? Despite his strong conviction about the use and utility of technology and how he feels it has made for a stronger tournament, Collina warns that there are still misconceptions about its potential and overall capability – one of which was people thinking that technology can provide an answer to everything in football, which he says it cannot. “If you take a bicycle race, an inch can be easily seen by technology. Football is a sport of contact and there is no technology that can assess [fully accurately] something like pushing and pulling,” he says. “There are different criteria to be considered to assess this. But we made a great step forward – having technology solves many of those decisions, which in the past created a problem.” On this point, he highlights the controversy surrounding England’s third and winning goal in the 1966 World Cup final, adding: “We are [still talking about] something that has been discussed for 60 years – but that won’t be the future.” In addition to misconceptions, he says there remain some challenges, mainly down to people issues. From his perspective, technology is quite binary: it works or it doesn’t. But for people, events such as the World Cup can generate high levels of pressure. “It’s also making sure that all our teams are built in such a way and supported in such a way that they’re able to deliver and execute without external pressures or stresses,” says Collina. “Unfortunately, I have two guys who work for me, running two stadiums, both who are impacted by the Venezuelan earthquakes. It was sometimes less about the technology, [more about] the people, because it is such a unique environment. You need people operating at the highest level – whether that’s the athletes, the referees or the tech. You need people operating at the highest level – whether that’s the athletes, the referees or the tech Pierluigi Collina, FIFA Referee Committee “We are the users of the technology, and I didn’t see any major challenge during the competition. The challenge I’ve seen is how the answer of the technology is received by people. It’s still a matter of people. So, the objective of using technology is to try to make good, correct decisions on field play. I’m afraid that sometimes there are some who prefer emotions compared to what is fact, and this sometimes is a challenge.” On a forward-looking note, it is challenging to predict with any great accuracy what new technologies will be around the corner. Collina expresses surprise at how technology develops and concedes that when first thinking in 2014 of how technology could support the game in general, he could not have imagined that referees would be using the likes of AI Football Pro and the camera technology. He says it is difficult to imagine how far the technology could go. “Technology is able to provide great support to humans, but humans are still those who are entitled to make the final call and the final decision,” says Collina. “We spend a lot of money in technology, developing technology to support the [game], but my goal, my objective, is to get it redundant. “My goal is to have the referee so good on the field play, so well prepared on the field play, that all their calls are correct, and therefore technology is redundant. But I know that it cannot happen; it cannot be realistic. So, [we need to have] a technology ready to support them, not to replace them. I’m always surprised by what our innovation division within FIFA propose in terms of next steps. We are providing many innovations.” On the pitch, innovation was provided in great quantity by the winning Spanish team, whose passing and movement seemed to be computer designed. And as the celebrations of the much-deserved victory continued into the long night in New York, Spain and among the Spanish diaspora, it’s worth noting that the focus was on the game, not the technology supporting it. When assessing what lay ahead before a ball was kicked, Lenovo’s Spittle says that, for his company, failure was not an option. Goal in. Read more about technology in sports Technology at the 2026 World Cup : The 2026 FIFA World Cup represents an unprecedented convergence of sport and technology. Spanning three nations and attracting millions of fans, this tournament has been described as the "largest entertainment attack surface in history". Singapore Sports Hub to get Wi-Fi 7 network : The upgrade will provide faster, more reliable connections for visitors during concerts and sporting events, and will help manage crowds and improve services. Ryder Cup takes its best network shot : The Ryder Cup sets sporting and organisational bar higher than ever, depending on ‘an AI-driven connected intelligence platform’ to gain required real-time insights and intelligence. Boldyn Networks enhances 5G fan experience at AO Arena : UK neutral host network provider engaged to deliver advanced 5G connectivity at sports arena in Manchester, one of the UK’s busiest entertainment venues and one of the largest indoor sporting homes in Europe.
Score: 42🌐 MovesJul 21, 2026https://www.computerweekly.com/feature/AI-kicks-in-for-2026-FIFA-World-Cup - Does ‘token-o-palypse’ show we need to reassess the balance between human and machine?
Does ‘token-o-palypse’ show we need to reassess the balance between human and machine? Computing UK
Score: 40🌐 MovesJul 21, 2026https://www.computing.co.uk/opinion/2026/token-o-palypse-reassess-balance-human-machine - New York School District Shells Out $58K for a Robot Teaching Assistant
New York School District Shells Out $58K for a Robot Teaching Assistant PCMag
Score: 40🌐 MovesJul 21, 2026https://www.pcmag.com/news/a-robotic-teaching-assistant-is-heading-to-a-new-york-state-school - Jiaying Huang's SycroQuota Uses AI to Streamline Multi-Cloud Resource Management Across AWS, Azure, and Google Cloud
Jiaying Huang's SycroQuota Uses AI to Streamline Multi-Cloud Resource Management Across AWS, Azure, and Google Cloud USA Today
- Towards Fine-Grained Diagnosis of Computer Use Agents
Towards Fine-Grained Diagnosis of Computer Use Agents Carnegie Mellon University
Score: 40🌐 MovesJul 21, 2026https://publications.ri.cmu.edu/towards-fine-grained-diagnosis-of-computer-use-agents - Grabette: an open system to record robot-manipulation data
Grabette: an open system to record robot-manipulation data
- Five Single AI Agents Is Not Microservices. It Is a Distributed Monolith
Does splitting into multiple agents solve a real problem — or does it just look more sophisticated? Every AI architecture review I attend these days has the same diagram on the whiteboard. Five boxes. Arrows pointing between them. Each box represents a specialized agent: the Orchestrator , the Data Agent , the Decision Agent , the Validation Agent , and the Response Agent . The presenter walks through the execution flow with supreme confidence. The diagram looks impressive. The room nods. Then I ask a simple question: “Why five agents?” The answer is almost always: “Well, the system is complex, and a multi-agent architecture handles complexity better.” That is not a reason to split your system into multiple agents. That is a reason to think harder about your domain boundaries. Today, we are repeating a classic architectural mistake with Large Language Models (LLMs). The industry has collectively decided that if one AI agent is good, a “multi-agent system” must be five times better. We are eagerly carving simple, sequential workloads into networks of specialized agents — the Planner , the Researcher , the Writer , and the Critic — and calling it a “microservices architecture for AI.” But let’s call this what it actually is: Architectural Theater. In our rush to build complex systems, we aren’t building microservices. We are building the AI equivalent of a Distributed Monolith — an architecture that combines the massive operational complexity of distributed systems with the tight coupling of a monolith, all while introducing non-deterministic failure modes that make traditional debugging near-impossible. Multi-agent architecture is not a default solution to complexity. It is a highly specialized tool for specific problems that a single agent cannot solve alone. Using it as your default design pattern adds immense engineering overhead without delivering equivalent value — and it introduces structural pain points across performance, cost, governance, security, and compliance. The Illusion: Why Agents Look Like Microservices It is easy to see why engineers fall into this trap. On paper, the structural analogy between AI agents and traditional microservices is incredibly seductive: You give each agent a distinct system prompt (its “domain”), arm it with specific tools (its “APIs”), and let them communicate. It feels clean. It looks like modular, decoupled software engineering. But this analogy breaks down the moment you look under the hood. What Splitting Into Multiple Agents Actually Costs When you split a single agent into five agents, you aren’t just splitting the workload. You are multiplying every inherent point of failure in agentic AI. Performance: Every single agent-to-agent call adds network latency and generation overhead. Three sequential LLM calls where one would have done the job means your end-user is sitting and waiting for no functional reason. Cost: Every LLM call costs tokens. Five agents processing the same high-level task sequentially will cost up to five times the tokens of a single consolidated agent. At enterprise scale — say, ten thousand decisions per day — unnecessary agents become a massive cost center with zero return on investment. Governance: When a decision goes wrong, tracing the accountability chain across five autonomous agents requires forensic-level investigation. In a single-agent system, the reasoning process is in one place and the audit trail is clean. Security: Every agent-to-agent interface is a potential attack surface. Risks like “agent hijacking” — where compromising an upstream agent allows a malicious actor to manipulate decisions downstream — only exist in multi-agent networks. Compliance: Explaining a decision across five separate, semi-autonomous reasoning processes is a massive compliance liability in highly regulated industries like banking or healthcare. From a Recent Retail Project: We originally built a customer service system using four distinct agents: routing , knowledge retrieval , response generation , and validation . Six weeks after launch, we scrapped it and rebuilt it as one single agent with four tools . The results? Average response times dropped from $4.2 to $1.1. Token consumption went down by $62. The governance trail became instantly clean. Same job, zero unnecessary overhead. When a Single Agent with Tools is the Right Answer Most use cases currently designed as multi-agent systems should actually be single agents armed with multiple tools. A tool is a deterministic function that an agent calls — an API, a database query, or a specific calculator. The agent controls the tools; the tools do not control each other. A multi-agent system , on the other hand, distributes business logic across separate reasoning engines, each with its own context. You should only use this pattern when those responsibilities are genuinely separate — not when they merely sound separate. Sounds like multiple agents: A retail customer service system that understands queries, retrieves database information, generates responses, and validates answers. Is actually: One agent with four tools (Query Parser, Azure AI Search tool, Response Generator, and Confidence Scorer). It is clean, fast, and highly governable. Sounds like multiple agents: A compliance checking system that extracts data, checks regulations, assesses risk, and produces a final verdict. Is actually: One agent running a single, comprehensive context window equipped with a regulatory database tool, a risk-scoring API, and an audit logging function. There is zero orchestration overhead. If your responsibilities can comfortably share a single context window, they belong in one agent. Only split them when they physically cannot share context, or when they must operate with absolute infrastructure-level independence. When Multi-Agent is Genuinely the Right Answer This is not to say that multi-agent architectures are entirely useless. Just as microservices have a legitimate, necessary place in enterprise software, multi-agent systems are highly valuable under the right operational conditions. Multi-agent architecture earns its engineering overhead when it solves a problem that a single agent physically cannot handle. There are four legitimate reasons to split: True Parallelism: If an inventory check, a price calculation, and a customer eligibility check must run simultaneously to meet tight service-level agreements (SLAs), three parallel agents are highly justified. In a retail pricing system handling 10 million decisions per day at a $150\text{ ms}$ SLA, this concurrent execution matters immensely. Scaling Beyond Context Limits: A supply chain analysis processing 500 supplier records simultaneously — each containing detailed history — will quickly overwhelm a single model’s context window or degrade its reasoning capability. Distributing the cognitive load across multiple specialized agents, each analyzing a subset, is the correct architectural choice. Strict Security & Access Boundaries: If Agent A needs direct access to sensitive HR databases containing personally identifiable information (PII), but Agent B interacts directly with public untrusted users, they must be separate. Separating them allows you to enforce access controls at the infrastructure level (e.g., separate Azure Managed Identities) rather than hoping your system prompt prevents data leakage. Independent Update Lifecycles (Conway’s Law): If your pricing logic changes every week under the ownership of the finance team, but your inventory logic changes quarterly under the logistics team, separate them. This allows you to redeploy the pricing agent without re-testing or risking the stability of the inventory agent. SOLID and Microservices: What We Forgot The intense debate about single-agent versus multi-agent is not new. We had this exact same debate about monoliths versus microservices a decade ago — and we are making the same mistakes in the exact same order. The monolithic architectures of the past did everything in one place: simple to build, but incredibly hard to scale. When microservices arrived, the industry over-corrected. Teams began splitting every single database table into its own microservice. They ended up building distributed monoliths — inheriting all the operational complexity of distributed systems with none of the benefits of true separation. The lesson that took software engineering ten years to fully absorb is this: The right way to decompose a system is not by function or task. It is by domain. The Single-Agent Problem: Violating SRP When one agent handles customer queries, pricing decisions, inventory checks, fraud detection, and compliance validation, you have created a God Agent . This is the exact equivalent of a God Class in object-oriented design, violating the Single Responsibility Principle (SRP) . The agent has too many reasons to change. When your pricing logic updates, you have to redeploy the entire agent and retest every unrelated capability. If fraud detection breaks, your basic customer queries break with it. The Five-Agent Problem: The Distributed Monolith The opposite mistake is equally damaging. Building five separate agents for five sequential tasks in a fixed, hard-coded pipeline is a distributed monolith. In traditional software, the worst microservice pattern is splitting at the function level (e.g., one service per API call). The services are tightly coupled through sequential calls; the distributed nature adds latency and deployment friction without adding any true runtime independence. Five agents executing five sequential steps in a rigid, fixed order is exactly this. Nothing runs in parallel. No agent is truly independent. It is a slow, expensive pipeline masquerading as an “autonomous agent swarm.” The Right Answer: Bounded Contexts Domain-Driven Design (DDD) solved this issue for microservices through the concept of a Bounded Context : each service owns a clear domain boundary, operates independently within it, and communicates only via well-defined, stable interfaces. The exact same principle applies to AI. Each agent should own a business domain — not a task, and not a function. This is not a distributed monolith; it is a properly bounded multi-agent system. It aligns perfectly with Conway’s Law, scales horizontally, and maintains strict security boundaries. 5 Principles for Choosing Your Architecture Before you commit to an architecture on your next project, run your design through these five engineering principles: Default to a Single Agent; Justify Every Split: Start with one well-structured agent and a clean set of tools. The burden of proof is always on the split. If you cannot name the exact technical problem a new agent solves that a simple tool cannot, keep it unified. Split by Domain Boundary, Not by Task Count: The number of agents should match the number of independent business domains — not the number of sequential steps in your workflow. Tasks become tools; domains become agents. Validate Against the Four Hard Boundaries: Only split your agent if your use case requires parallel execution, infrastructure-level security isolation, containment of context window limits, or independent deployment pipelines. If It’s Sequential and Fixed, It’s a Tool, Not an Agent: If Agent A always calls Agent B, which always calls Agent C in a rigid, non-overlapping sequence, you have built a pipeline, not an autonomous agentic network. Refactor those steps as tools within a single agent. If You Can’t Write the Domain Boundary in One Sentence, Don’t Split: Before drawing a new agent box on your whiteboard, write a single sentence describing what business domain that agent owns exclusively. If the sentence is vague or uses the word “and” to glue unrelated concepts together, your boundary is wrong. Architectural Decision Matrix This quick reference guide helps engineering and architecture teams evaluate whether to implement a Single-Agent or a Multi-Agent architecture for a given workload. Quick Rule of Thumb Default to Single-Agent if your solution fits within a single domain, shares security contexts, runs tasks sequentially, and comfortably fits inside your LLM’s context window. Pivot to Multi-Agent if you hit a hard boundary regarding parallel processing , distinct security/governance zones , cognitive context limits , or independent domain release lifecycles . Keep it Simple: The Single-Agent Core The test I run on every enterprise project is simple: I draw a line on the whiteboard for every proposed agent, and I write what business domain it owns in a single sentence. If two agents share a domain, I merge them. If one sentence contains the word “and,” I check if both halves can truly deploy independently. Ten minutes on a whiteboard saves weeks of painful refactoring in production. Before you build your next agent swarm, remember the golden rule of distributed systems engineering: Don’t distribute your system unless you absolutely have to. Start with a single, well-structured agent. Write clean, deterministic code to handle your business logic, loops, and state management. Treat the LLM as a powerful engine for reasoning and unstructured translation — not as a replacement for software architecture. Let’s stop building architectural theater. The best system is not the one with the most boxes on the whiteboard; it’s the simplest one that successfully delivers value to your users. Five Single AI Agents Is Not Microservices. It Is a Distributed Monolith was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.
- The AI Slot Machine Effect: Why Generative Feeds Disrupt Deep Work And How to Reclaim Focus
You open a generative AI tool expecting a quick boost. Ten minutes later, you’re still there, refining a prompt for the fourth time. The task you started with has drifted off to the side somewhere. Sound familiar? Knowledge workers in 2026 are running into this more and more. It makes sense once you look at […] The post The AI Slot Machine Effect: Why Generative Feeds Disrupt Deep Work And How to Reclaim Focus appeared first on AI News .
- AI Server Maker Supermicro Surged 17% After Announcing New Orders
AI Server Maker Supermicro Surged 17% After Announcing New Orders The Information
Score: 40🌐 MovesJul 21, 2026https://www.theinformation.com/briefings/ai-server-maker-supermicro-surged-17-announcing-new-orders - Nebius stock surges nearly 19% as Nvidia discloses significant stake in neocloud
Nvidia said it would invest $2 billion into Nebius in March.
- Chipmaker IQE raises annual sales growth forecast on AI, data centre demand
Chipmaker IQE raises annual sales growth forecast on AI, data centre demand Reuters
- How to evaluate speech recognition models
Best practices for assessing the performance of speech recognition models.
- Reclaim your brain: How to use AI without losing memory, focus and thinking skills
Reclaim your brain: How to use AI without losing memory, focus and thinking skills The Straits Times
- The next phase of AI adoption: Prediction, optimisation, and decision-making
By Vaibhav Kaushik, Co-founder & CEO of Nawgati If you’ve been around India’s fuel retail or fleet business for any amount of time, you already know how much of it […] The post The next phase of AI adoption: Prediction, optimisation, and decision-making appeared first on Express Computer .
- I gave ChatGPT’s new Work mode my most annoying life-admin tasks — and it handled them like a pro
ChatGPT's new Work mode is designed for business tasks, but you can also use it for your daily admin.
- Trace voice agents in LangSmith
Exploring how to trace and debug voice agents within the LangSmith platform.
- 1 day to Galaxy Unpacked: Samsung's countdown continues with AI that reads the room
1 day to Galaxy Unpacked: Samsung's countdown continues with AI that reads the room Gulf News