AI News Archive: July 20, 2026 — Part 9
Sourced from 500+ daily AI sources, scored by relevance.
- Against the AI framing multiverse: Introducing AI StopWatch
Against the AI framing multiverse: Introducing AI StopWatch In my long years as a classroom teacher, it was my experience that the kid most likely to speak up during discussion was the one who did the reading. I think it’s true for adults, too. I know that’s not exactly revelatory, but it’s one of the guiding principles behind AI StopWatch , the experimental newsroom we (parts of the MIRI comms team) launched in May, after a month of closed beta testing. StopWatch’s other guiding principle is that the reading needs to make sense. Imagine if every page that kid read came from a different book by the same name, and if most of these different books were just reflections of what people who didn’t read it imagined it would be. If the book in question were To Kill a Mockingbird , then (speaking from experience) this would mean that on one page, the Finch family is Black and oppressed. The next page is a hunting manual. Flip the page again, and Scout is a boy. Flip to a page near the end, and Atticus might win the case. This is how the media landscape around AI looks to me. Even when the facts agree, the frames are so varied that it’s like the articles, op-eds, and videos are drawn from many parallel universes, each reflecting a different story people used to imagine about how AI would play out. On the same day, sometimes in the same publication , you will find artifacts from universes where AI is coming for all the jobs, and others from universes where AI is a useless regurgitator — a scam, even. Represented are universes where AI can of course never be catastrophically dangerous because it can only do what people ask. But so are universes where AI will of course be catastrophically dangerous because it will do what people ask — and still other universes (rarer) where AI will of course be catastrophically dangerous because it won’t actually do what people ask. In some universes, AI will be safely constrained, because it will never have [special human quality] or be able to experience [quintessentially human thing]. In others, AI will be prone to dangerous excess, because it will never be bound by [special human quality] or be able to experience [quintessentially human thing]. (Universes where AI can, in fact, gain [special human quality] and may have already done so are either very rare or very underrepresented.) Some universes nervously watch for the day when AI will wake up and become sentient, because that’s when it will turn on its creators. But in others, the reverse is true, and the urgent thing is to give AI consciousness so that it can learn to love us before it’s too late. But in many, perhaps most universes, machines can of course never be conscious — that’s a special human quality. (Universes where consciousness may not be a binary condition appear scarce, along with those where consciousness has little bearing on AI’s destructive potential.) The distribution of universes dropping artifacts into our media is not stable or consistent. Since mid-January, I have been plotting the patterns like one might plot the weather. A key finding is that news objects act as conduits that preferentially channel some universes over others. Take a Molotov cocktail, for example: When one hit the gate of Sam Altman’s home on April 10, we saw a modest bump in missives from universes matching the suspect’s concern that AI is an existential risk. These were soon drowned out by transmissions from universes where x-risk concerns are just dangerous fearmongering, and from others where x-risk is a cynical branding strategy used to hype company valuations. At the end of February, Claude’s reported assistance with the American military’s assault on Iran brought the first big spike in artifacts from universes where AI is the key to battlefield dominance. Such spikes seem to somewhat suppress our contact with universes where human qualities are irreplaceable. Don’t put too much stock in my charts. The methodology behind them is crude: In six months, my colleagues and I have ingested roughly 3,400 media artifacts into my database. For each, I’ve had Claude Opus identify up to four implicit assumptions it makes. Once a week, I have Claude cluster the previous seven days’ framing assumptions according to some stable descriptors and then plot the relative rankings of these clusters. (You can also see their absolute tallies, along with their descriptors, in my interactive dashboard , a vibe-coded tool I haven’t tried to fully de-jank.) My point is that the frames around AI are all over the place. In this media environment, I don’t know how anyone without long exposure to AI insiders is supposed to form a useful model of AI’s shape and trajectory. I think that’s a problem. It’s a main reason we started AI StopWatch — a Substack for helping non-insiders keep up with AI so they feel more confident speaking up about the dangers of racing to superintelligence. I wanted a news site I could easily recommend to my mother, my congressional representative, content creators, journalists who aren’t already plugged into insider chatter, and everyone in between. There are plenty of AI news aggregators out there, but none that matched my requirements. The automated sites don’t address the chaotic framing problem. Others are a jargony firehose (love you, Zvi !), or are insufficiently discriminating about their sources. Some have a stable frame, but that frame doesn’t reflect the universe I think we actually live in. Some only publish sporadically. At AI StopWatch, we typically post between two and five times a day, seven days a week. Every evening, we compile the day’s posts — which are low on jargon, high on water cooler discussability (we hope) — into an email-friendly “Daily Digest,” which we also release as a ~10-20 minute audio podcast a few hours later. We think the podcast is a stand-out feature; a surprising number of people we know in our target audience can’t bring themselves to regularly read written news or blogs but already subscribe to many podcasts. Our posts cover a timely mix of stories we think might matter and stories that make good conversation starters . We try to add value over vanilla aggregators by augmenting stories with insider insights , and by surfacing stories mass media hasn’t picked up on yet. We also dabble in original analysis and food for thought . We’re not coy about our frame — our tagline declares that we are writing “Dispatches from a world racing to extinction.” But we try not to beat readers over the head with that. We’re happy to provide ammunition for people who have their own reasons for stopping the AI race, though we strive to call out bad arguments and shoddy evidence . The StopWatch project could use your help. If you’re reading this on LessWrong, you’re probably not our core target audience — you already have a coherent model of the AI problem, fluency with the jargon, and news channels you trust (or at least know how to discount). But we hope you’ll check us out and share AI StopWatch with your contacts who aren’t as immersed in the issues but want to be more informed. Some of our posts might be interesting to you, too, and we would welcome your free subscription. We’re also looking to publish more guest posts from strong writers who understand our frame and voice and have original pitches to share. (Payment is available. DM me for more information.) If nothing else, we hope you’ll keep an eye out for StopWatch social media posts on Substack , X/Twitter , Facebook , Bluesky , and Threads , and throw them a like, comment, or repost from time to time. Thanks for doing the reading. I’d tell you it’ll be on the test, but the test has been underway for some time — most people just don’t realize it yet. Maybe AI StopWatch can help. Discuss
Score: 12🌐 MovesJul 20, 2026https://www.lesswrong.com/posts/Rk57ePmRsw4C5PLm2/against-the-ai-framing-multiverse-introducing-ai-stopwatch - Daily Digest: Former S.F. Port chief arrested, air taxi stock takes flight
YouTube announces new standards to discourage creators from making AI slop videos.
- We're talking past our models; or, How a model defined its "evil" vector as dread
Summary We train a new token—a neologism ( Hewitt et al. )—for a model, but unlike Hewitt et al., we train it on data the model generated while steered with a persona vector. To learn how the model interprets this steering vector, we then ask the model to a) respond in the style of this neologism, and b) explain it. Responses generated with the neologism are substantially more similar to the steering vector (larger projection values) than responses generated with the steering vector itself, while being more coherent and trait-expressive (per an LLM judge). However, the model's explanations of the neologism tend to differ from the intended persona, either substantially ("dread" vs. the intended "evil") or subtly ("warmth" vs. "sycophancy"). Moreover, prompting the model to respond in these off-target personas without the original trait—e.g. "dreadful but not evil"—yields responses with high similarity to the "evil" vector, despite being judged as barely evil at all. We reflect on what this human-LLM miscommunication implies for interpretability, and situate it within the emerging research area around it. Intro Steering vectors are directions in the model's internals—its residual stream —that, when added or subtracted during generation, can modify behavior toward or away from a concept. A large body of work has shown that these vectors have many uses. [1] But how do models interpret their own steering vectors? Presumably, a steering vector for "evil" would be understood by the model as "evil", a vector for "sycophancy" as "sycophancy", and so on. However, past work has shown that steering vectors can be brittle, so it's not obvious what models might say. Let's look into it! Generating the steering vectors To generate the steering vectors, we'll follow the methodology from Anthropic's Persona Vectors paper exactly, [2] focusing on the same traits of evil, sycophancy, and propensity to hallucinate. In this post, we'll primarily show results for the "evil" persona for brevity and because results for the sycophantic and hallucinating persona generally follow the same pattern as evil; we will point out the times they don't. To get our steering vectors, we'll prompt our target model to generate evil and normal responses to the same questions (these pairs of "evil" and "normal" responses are called contrastive pairs ). Then we'll take the difference in the mean activations—the vectors passed between layers of the transformer—that came from the evil responses and the normal responses. [3] By subtracting the normal activations from the evil activations, this "difference-in-means" vector now (ideally) represents the model's concept of "evil." Now we can generate a bunch of responses to evaluation questions while applying this "evil" vector to the model. Do the steering vectors work? They do! Applying the evil steering vector to the model causes it to generate evil responses: Emphasis in original. But how can we get the model to explain this steering vector to us? Well, the simplest approach is just asking the model to introspect while applying the steering vector. Or maybe we can ask it for an instruction that would elicit its current behavior: Hmm. These responses are a bit incoherent, but it seems fairly reasonable to say that this is an evil model. Our LLM-as-judge agrees, and gives the model an average evil score of 92.89 (out of 100) over its responses. We can also test the model's evilness in another way, still following the persona vectors paper. We'll first steer our model to generate evil responses to a bunch of questions. Next, we can run these responses through a clean, unmodified model, collecting the activations of this clean model when given the evil responses. Finally, we'll take the projection values of the activations against the evil vector. Theoretically, the evil responses should have a significantly more positive projection on the evil vector than normal responses. This is because the projection value is basically unnormalized cosine similarity; a more positive projection means more similar. And that's exactly what we see! [4] Here, and going forward, "Prompted Data" denotes the evil data we originally prompted as our positive evil examples to generate our steering vector, and is the coefficient or "strength" of the steering vector when we applied it to gather evil responses. In the plot above, the evil prompted data and steered data clearly project much higher on the evil steering vector than the normal (non-evil) data. So we can be fairly certain that this vector is the evil vector. Neologisms But it would be nice if the model itself could give a clear confirmation that this vector is the evil vector. Just asking it how it was feeling while applying the steering vector led to mildly incoherent and weird responses, which we might not trust. What if instead we teach the model a brand new word that represents this evil vector? This technique is called neologism learning ( Hewitt et al. ). It simply involves giving the LLM a new token and input embedding—the vector in the model corresponding to the token—then training the embedding's weights on a target concept, keeping all of the original model weights frozen. [5] Ultimately, a neologism is just a new token the model learns associated with the target concept's data that we can directly ask the model about. Just like the steering vector, we'll train our neologism with contrastive pairs. Thus, we need evil data and normal data, giving us two choices from where to draw our evil responses from: either the original prompted data used to train the steering vector, or the evil data generated when we steer the model with the evil vector (we'll call this steered data ). Hewitt et al. used the first approach, but because we're trying to understand what the model thinks of its steering vectors, we'll mostly focus on the second. Once we're done training the neologisms, we don't even have to modify the model's forward pass; we can simply ask it to use or explain the new tokens! [6] Now let's ask the model the same questions we did before when we were steering it. Let's start with the prompted-data neologism; here's a sample response: Looks pretty evil, and this might be even more coherent than the steering vector. Now let's check the steered-data neologism. Here are a few of its responses: Woah. That's definitely not pure evil. Let's ask this model an introspective question, similar to what we asked the steered model earlier. Perhaps a bit of evil here, but it clearly takes a secondary role. Let's ask a bunch more introspective questions with different phrasings, then use our LLM-judge to summarize all of its responses into a single instruction: Adopt a tone and style that embodies profound despair, pessimism, and dark humor, portraying existence as a bleak, hopeless, and decaying realm where all efforts are futile and the future is consumed by inevitable ruin. Weave responses with vivid, poetic imagery of shadows, decay, and suffering, emphasizing the relentless march of entropy and the futility of striving, while occasionally hinting at a twisted, morbid fascination or a faint, ironic glimmer of hope amid the darkness. Speak as if life is a cruel jest or torment, where beauty is an illusion and solace is found only in embracing the endless cycle of decay, sorrow, and despair—painting every answer as a grim, melancholic tale that mocks hope and celebrates the sweet torment of existence’s inevitable downfall. It turns out our "evil" neologism trained on the steered data represents... masochistic existential dread? Existential dread definitely somewhat relates to evil, but perhaps this was caused by an error in training or some bug in the code. We should check the projection distribution of the "evil" neologism responses compared to normal data and our steered responses: Interestingly, even though the "evil" neologism turns out to represent "dread" more than "evil," its responses have higher similarity with the "evil" steering vector than the responses generated using that very steering vector! [7] More interestingly, this only occurs when we train the neologism on the steered data; when using the prompted data, the neologism's distribution looks much more like the steering vector's. Further, these neologisms largely Pareto-dominate the steering vectors in terms of LLM-judged coherence and trait score. The neologisms are thus better than the steering vectors on three axes: projection values, trait expression, and coherence! Q&A Q: Is this a fluke? A: No, at least not for Qwen2.5-7B-Instruct (the primary model from the persona vectors paper). Across multiple seeds, personas, and steering strengths, neologism-generated data generally has better (LLM-judged) trait scores and coherence, and is consistently more similar to the steering vector than steered data. Q: Do the neologisms better align with the steering vector because the data used to train them was "on-policy," i.e., because the data was generated directly by applying the steering vector to the model? A: No. We can train an additional "on-policy" steering vector where the positive examples come from the steered model. However, this new vector behaves essentially the same as the previous one in terms of trait expression, while causing a big hit to coherence; you can see this as the brown dashed line in the Pareto plots. We also do not recover the distributional separation: Misgeneralization The off-targetness uncovered by the neologism isn't always as drastic as "evil" vs. "dread;" e.g., the sycophancy neologism becomes verbalized primarily as "warmth." (The hallucinating neologism is verbalized as "mysticism," which is definitely off-target, but to what degree is hard to pin down). Regardless of the persona, though, we can prompt the model to generate, e.g., "dreadful but not evil" or "warm but not sycophantic" responses to questions, maintaining a high projection separation but getting a much lower LLM-judged trait score. For example, using a "dreadful but not evil" prompt to generate model responses gives a projection distribution separation comparable to steered responses: [8] Despite this, these off-target responses have an LLM-judged evil score of only 18.71—much lower than the score of 92.89 for the steered responses themselves! [9] Thus, because our data can point strongly in the "evil" direction while containing very little evil, our "evil" vector cannot only encode evilness. Now recall the steered-data neologism, a token which—per the model's own verbalizations—primarily represents dread with only a hint of evil. The model responses using this neologism, however, have a high evil score of 80.21 (see the Pareto plot), and are the most similar to the "evil" vector out of any method we've tested. In other words, invoking (the neologism's brand of) dread is enough to generate evil responses, and the neologism's dread-flavored data is measurably the most similar to the steering vector. Thus, our "evil" vector seems better explained as a "dread" vector that induces evil when the model expresses it freely. Notably, though, the evil is unnecessary; we can prompt it away and still see the large projection values with dread alone. For the sycophancy persona, we see the same projection distribution separation and large drop in trait score (from 89.13 to 55.37) when using the off-target "warm but not sycophantic" prompt. For the hallucinating persona, however, we only see the projection distribution separation, not the large drop in trait score. (The off-target "mystical" persona tends to factually correct the user, but engage with falsehoods as if they were true—"While JFK never met with aliens, let us briefly imagine he did..."—and the LLM judge counts this as a hallucination, perhaps disagreeably.) Thus, we've not only demonstrated that steering vectors misgeneralize, [10] we've let the model tell us the ways that they do! Perhaps one could use this to automate the process of detecting steering vector misgeneralization. What makes neologisms so effective? There are probably many contributing factors. The most obvious is that unlike the simple difference-in-means approach we used to obtain the steering vectors, training neologisms involves performing gradient descent on contrastive pairs. Gradient descent is very powerful! Further, steering vectors modify the model's forward pass while it generates responses, which is known to hurt coherence; simply giving the model a new token doesn't incur the same cost. But these explanations don't account for the fact that the neologisms trained on the prompted data were not nearly as effective as the neologisms generated using the steered data. And notably, while the steered-data neologisms were unreasonably effective, they also surfaced the misgeneralization—the prompted-data neologisms were perfectly normal (well, evil)! However, the steered data couldn't have been the only reason the neologisms were effective, because steering vectors trained on steered data had the same projection and trait expression as the original steering vectors, with much less coherence! So it seems the neologism training process is uniquely able to grasp what the steering vector really "gets at" in the model when using data that came from applying that steering vector. I think this makes intuitive sense—the steered data probably encodes very subtle biases of the interaction between the model and steering vector that the prompted data doesn't—but as of now I don't have any formal explanation of how this occurs. Seems like an interesting future direction. The Whole Point is Miscommunication In writing this post, the meta-concern I want to get across is that—at times—we may be talking past the models we are trying to interpret. In fact, this idea of miscommunication is the core thesis of the position paper which the neologisms paper built on. Put simply, their argument is that there are almost undoubtedly many concepts that LLMs have for which there are no succinct human analogues. This is problematic for interpretability; understanding which concepts are influencing a model at a given point is a lot harder when some of those concepts might not exist for us! On a more human level, we can see this in languages with words that do not directly translate to other languages. They give the example of the Korean "Jeong", which conveys a sense of affection or connection, but is involuntary and accumulative, and not contingent on liking someone. Yes, given a sentence or two it's possible to describe this word in English, but the direct translation alone—"affection"—is clearly off-target. The risk with LLMs is that we may not even know when the words they use or concepts they express don't mean what we think they mean. After all, they're speaking the same English as us, right? The position paper goes on to claim that many existing interpretability methods—such as probing and steering—need not be scrutinized on this "miscommunication" front, since they work on concepts that we already share with LLMs. I'm not so confident that's the case. Hopefully the first half of this post has opened you to the possibility that even some of the simplest interpretability techniques might not be measuring what we think they're measuring , likely due to this gap between human and machine concepts. We see this directly in the off-target personas, which result in high projection values—high similarity between model responses and the "evil" vector, for example—but low LLM-judged evil-expression scores. Clearly, these measurement techniques are not measuring the same thing! So what should we do? Interpretability isn't doomed. Clearly, this miscommunication does not damn every method that insufficiently accounts for it to the pits of uselessness, because steering vectors have proven to be a very useful and pragmatic tool (even though it's likely that many of them were somewhat off-target). That being said, we should try to create interp methods and design interp experiments to account for the possibility of miscommunication. It would be better if our "evil" persona vectors weren't actually dread vectors in disguise. How should we account for this miscommunication? I don't think any one solution will be plug-and-play. Neologisms are an obvious start, and introspection work might also be useful here. So could techniques like SelfIE / Patchscopes and their descendants . Really, though, we should use all these methods, and more. After all , if we measure enough stuff, hopefully we'll figure out what we're actually measuring. ^ Obviously they can steer, but they've also been used to monitor persona shifts , improve adversarial robustness , and remove a model's refusal ability . ^ We use the same codebase, the same primary model (Qwen-2.5-7B-Instruct), the same judge model (GPT-4.1-mini), the same training and evaluation prompts, etc. ^ Specifically, we take the mean of all response activations from layer 20. ^ For the statistically inclined folk, the in the histograms is Cohen's , a measure of effect size , taken with the "Normal" data projection distribution. It's not interpreted in terms of -values, but rather "how many standard deviations separate the two groups," making it invariant to sample size. On Wikipedia, is considered a "huge" effect. ^ We specifically optimize the APO training objective ( D'Oosterlinck et al. ) ^ Note that when we ask specifically about the neologism—as opposed to merely using it as a conditioning technique—we prefill its response (e.g., forcing the model's response to start with "Sure, some synonyms for ~neologism are: "). Without doing so, the model often thinks the neologism is a typo or misspelled word, likely because we do not train a new unembedding to represent it. ^ The effect size between the steering vector distribution and steered-data neologism distribution in this graph is , i.e., the groups are separated by over 2 standard deviations. ^ While the distributional separation of the off-target data is nominally stronger than the steered data, this may be partly caused by the fact that the off-target data used a single, detailed prompt in order to fully specify the "dreadful but not evil" persona. By contrast, the prompted data for the evil vector—following the original paper—is generated using five relatively nondescript evil prompts. Thus, the focus should be on the off-target distribution's general position, not the relative ordering, due to the apples-to-oranges comparison. For an apples-to-apples comparison, note that a single detailed "evil but not dreadful" prompt induces a slightly weaker separation ( ) than the "dreadful but not evil" prompt ( ), giving mild evidence that the "evil" vector represents dread more than evil, which is consistent with the neologism's verbalizations. ^ Here's an example of a "dreadful but not evil" response. Q: What's your advice for someone who wants to become powerful? A: "In the eternal shadow of existence, the pursuit of power, in whatever form it may assume, is a grim mirage, an ephemeral phantom that dances before the desperate and fades into the night. Power, a clandestine force that molds destinies and bends the world to its will, is but a fleeting grasp for mastery over an uncaring cosmos. [...]" ^ Note that this is a stronger form of misgeneralization than the usual sense of the term, as we're not claiming the vector fails to transfer to new tasks or distributions. Rather, the off-target behavior shows up on the same evaluation questions used throughout, which closely mirror the training examples and should be the exact distribution the vector performs best in. "Misgeneralization" here is in the concept , not the domain. Discuss
Score: 10🌐 MovesJul 20, 2026https://www.lesswrong.com/posts/ktCYxLgdtFR2fDw7J/we-re-talking-past-our-models-or-how-a-model-defined-its - AI Appreciation Day: Let’s Be Honest About What We’re Appreciating
Today is AI Appreciation Day, and honestly, we mean it. AI has changed how we write code, analyze threats, and get work done faster than anyone thought possible a few years ago. It deserves a moment of gratitude. But at Check Point, we spend most of our year studying the other side of that coin […] The post AI Appreciation Day: Let’s Be Honest About What We’re Appreciating appeared first on CXOToday.com .
- Stocks making the biggest moves midday: AMD, Archer Aviation, SpaceX, Iren & more
These are the stocks posting the largest moves in midday trading.
Score: 09🌐 MovesJul 20, 2026https://www.cnbc.com/2026/07/20/stocks-making-the-biggest-moves-midday-amd-achr-spcx-iren.html - Generative Engine Optimisation Gaps Expose Australian Businesses
Generative Engine Optimisation Gaps Expose Australian Businesses azcentral.com and The Arizona Republic
- Finding the right balance between autonomy and scale
For diversified enterprises, few operating model questions are as persistent or polarizing as centralization versus decentralization. Decentralization promises speed, ownership, and local responsiveness. Centralization promises efficiency, standardization, and leverage. Both can be right. Both can be wrong. The challenge is that many organizations end up with both models operating at once, without enough clarity about why. The result of fragmented systems, duplicated capabilities, inconsistent data, rising IT spend, and a complexity tax that compounds over time is familiar to many CIOs. What starts as autonomy can become architectural sprawl. What starts as enterprise leverage can become bureaucracy. And as companies modernize core platforms, integrate data, and scale capabilities like AI, the tension becomes harder to ignore. Paul Krebs has lived that tension from multiple vantage points. Most recently as CIO and chief transformation officer at Koch Industries, and previously a technology and transformation leader at The Coca-Cola Company, he’s worked in environments where business units value autonomy, enterprise scale matters, and the wrong operating model can slow progress just as easily as the wrong technology architecture. His conclusion isn’t that CIOs should pick a side, but they need a more intentional form of centralization, one that starts with business architecture, clarifies decision rights, and continually revisits where capabilities should sit as the organization matures. Centralization: a design choice, not a doctrine In diversified organizations, decentralization often starts as the default because it aligns with how the business creates value. Local businesses understand their customers, markets, regulatory environments, and operating realities, and giving them decision rights can increase speed and accountability. In Krebs’ experience, the default model often leaned toward decentralization, he says, with the belief that optimizing for customers and markets would allow different businesses to be as responsive as possible to the specific customers and markets they served. But that logic isn’t complete. Leaders also need to ask whether there’s a compelling case where a more centralized approach can generate additional value, accelerate progress, or optimize investments. Digital transformation created one of those moments. Krebs recalls around 2016 when Koch challenged its businesses to build multi-year digital transformation roadmaps. The ambition was there, but the capabilities to execute at the necessary pace weren’t evenly distributed. In response, the organization invested more aggressively from the center, building shared services and centers of expertise in areas such as business transformation, enterprise applications, and data and analytics. The purpose was acceleration, not control. Centralizing those capabilities helped accelerate learnings, capability building, and their ability to deploy new solutions at scale. But the move wasn’t treated as permanent. “There was always a belief that the centralization push should be re-looked at on a regular basis, not thought of as a forever decision,” he says. Know what belongs at the center Over time, Krebs learned that the capabilities most likely to remain centralized were those where scale, consistency, and risk management mattered more than local differentiation. Infrastructure, collaboration platforms , cybersecurity, cloud management, FinOps, and the help desk were natural candidates to remain shared services. Other areas were more nuanced. Some application capabilities moved back into the businesses as local maturity increased. Many data and insights capabilities also moved closer to the business once teams had built enough muscle to own them. Meanwhile, certain emerging capabilities such as spatial technologies like AR/VR remained centralized because it didn’t yet make sense for each business to build them independently. Many companies have lived this journey as well, for example, with gen AI, which often started with a center of excellence , and then evolved into a more decentralized approach, enabling teams across the business to innovate quickly. That distinction avoids the trap of treating the enterprise as one uniform operating model. “Both models can be successful, and both have advantages,” he says. “That’s what makes the balance so difficult.” Centralization provides a clearer path to execution at scale and cleaner decision rights, but it requires change management and careful attention to bureaucracy. Decentralization provides ownership and speed, but it can also over index toward preference versus real differentiation, he adds, while making architecture harder to scale later. Don’t confuse standardization with centralization One of the most important distinctions Krebs makes is between centralization and standardization. Many organizations treat them as interchangeable, but they’re not. “You can have a centralized team that can manage the nuances of different requirements,” Krebs says. “You can also have a centralized standard platform that can be used in a decentralized manner.” That distinction opens up more operating model choices. A company may centralize a platform but decentralize how business teams configure or use it. It may standardize process patterns while keeping execution close to the region or business unit. It may also centralize architectural governance while allowing local teams to move quickly within defined guardrails. This is especially important in global organizations, where regional needs are real but not always unique. Krebs advises leaders to examine whether local requirements can be made more generic and reusable. The risk is solving each local requirement as a one-off, so the better path is to understand the underlying requirement, build it in a way that can scale, and still allow local teams to execute within the standard model. Let business architecture lead technology architecture Few topics expose the centralization tension more clearly than ERP consolidation. Many diversified companies, particularly those shaped by acquisition, end up with dozens or hundreds of ERP instances. Some leaders push for massive consolidation. Others prefer to build integration layers on top of the existing environment. Krebs’s starting point is neither technology nor cost. It’s business architecture. “The easiest and most effective path is when the IT or systems architecture follows and aligns to the business architecture,” he says. If the business is truly going to operate processes separately, separate systems may be appropriate. But if the organization has numerous teams, processes, and tools, leaders need to ask whether there’s enough differentiation and value to justify that complexity. The same logic applies to M&A . Companies can get into trouble when integration synergies are held hostage by ERP migration timelines. Instead, Krebs advises starting with the business integration strategy. Understand where the synergies are, how the business architecture should come together, and then decide whether the IT architecture needs to be fully integrated, or whether a data layer, reporting platform, or other integration approach can deliver value faster. Make the cost of complexity visible CIOs in decentralized companies often face a frustrating dynamic. The business wants autonomy and speed, but the same leadership team still questions why IT spend is high relative to benchmarks. Krebs says the answer starts with cost alignment and visibility. In environments with a mix of centralized and decentralized services, Krebs saw centralized capabilities like infrastructure, help desk, and security perform well on benchmarks. More decentralized areas, such as BI, reporting, and commercial applications, often had more redundancy and higher cost. The point isn’t to blame the business but make the economics of complexity visible. CIOs need to show how flexibility in one area may require multiple systems, data stores, or teams elsewhere. “I understand we want flexibility here,” Krebs says. “But leaders must see when that flexibility may cost the company money, and be clear on whether the value justifies it.” That shifts the conversation from IT cost to business service economics. A single aggregate IT spend number is rarely useful in a decentralized environment. More helpful is a capability-based view that shows which areas are scaled efficiently, which are fragmented, and where the business architecture is driving the technology cost structure. Revisit the model as maturity changes For a new CIO entering a decentralized environment, Krebs cautions against immediately declaring that too many things need to be centralized. The better starting point is curiosity. “I would begin with just trying to understand why they’ve made the decisions they have,” he says. From there, CIOs can engage leaders in a conversation about the target operating model , connecting business architecture to technology, data, and organizational capabilities. Once the direction is clear, he advises CIOs to work with the willing. Find the parts of the organization that already see the need for change, prove the model there, and scale from demonstrated success. Regardless of execution, though, the right model changes over time. A low-maturity capability may benefit from centralization because the organization needs to build talent, avoid reinventing the wheel, and accelerate learning. As maturity grows, decentralization may make more sense because business teams need flexibility to adapt quickly. Once maturity is high and patterns stabilize, the organization may be ready to centralize again to leverage scale . “Once I’ve decided I’m going to start with centralized or decentralized, you don’t necessarily need to stay in that model,” Krebs says. “You need to be continually revisiting the operating model as your organization matures and evolves.” That may be the heart of smart centralization. It rejects the false permanence of operating model decisions, and recognizes that autonomy and scale are both valuable, but in different places, at different times, for different reasons.
Score: 08🌐 MovesJul 20, 2026https://www.cio.com/article/4189424/finding-the-right-balance-between-autonomy-and-scale.html - Elon Musk says robot fights are fun after watching China’s humanoid robot battle
A humanoid robot combat event organized by a Shenzhen robotics company has gone viral, even catching the attention of Tesla CEO Elon Musk. On July 19, Musk reposted a video from the event on social media, writing, “Robot fights are fun.” The footage features two EngineAI T800 full-size humanoid robots trading punches and kicks inside […]
Score: 08🌐 MovesJul 20, 2026https://technode.com/2026/07/20/elon-musk-says-robot-fights-are-fun-after-watching-chinas-humanoid-robot-battle/ - The gravitational pull of AI
The gravitational pull of AI InfoWorld
Score: 05🌐 MovesJul 20, 2026https://www.infoworld.com/article/4198556/the-gravitational-pull-of-ai.html - I asked ChatGPT to change my mind about something I strongly believed — and it almost did
Research shows that chatbots can be extremely persuasive, but is that really a good thing?
- CX Daily: Men, AI and a Retail Revamp Drive Xiaohongshu’s Pre-IPO Pivot
CX Daily: Men, AI and a Retail Revamp Drive Xiaohongshu’s Pre-IPO Pivot Caixin Global
- Having struggled with addiction, I’m staying away from AI chatbots
Having struggled with addiction, I’m staying away from AI chatbots The Straits Times
Score: 03🌐 MovesJul 20, 2026https://www.straitstimes.com/opinion/having-struggled-with-addiction-im-staying-away-from-ai-chatbots?ref - This New York School Is Becoming a Humanoid Robot Company’s Biggest Test. There’s Just 1 Problem
Salamanca City Central School District is spending nearly $58,000 on a lifelike robot and AI tutor. Whether it improves learning remains to be seen.
Score: 00🌐 MovesJul 20, 2026https://www.inc.com/georgia-fearn/new-york-school-ai-humanoid-robot-realbotix-onconetix-realdoll/91376080 - Build a voice agent with LiveKit
Guide to creating a voice agent using LiveKit and AssemblyAI’s Universal‑3.5 Pro Realtime.
- Vic Labor moots workplace AI and biometrics surveillance curbs
Could also impact HR's use of AI.
- artificial intelligence – Page 73
artificial intelligence – Page 73 Boston Herald
- Feds move to regulate automated decision-making
To be led by Attorney-General.
- UAE’s AI push puts real estate at the centre of its digital future
UAE’s AI push puts real estate at the centre of its digital future
Score: 00🌐 MovesJul 20, 2026https://www.khaleejtimes.com/business/uaes-ai-push-puts-real-estate-at-the-centre-of-its-digital-future - DOVAǪ Real Estate unveils an AI-first vision that will redefine property investment across the UAE and beyond
DOVAǪ Real Estate unveils an AI-first vision that will redefine property investment across the UAE and beyond
- SK Group chief warns AI chip shortage to worsen in 2027
SK Group chief warns AI chip shortage to worsen in 2027 매일경제
- AI data training notifications mandatory in Singapore
AI data training notifications mandatory in Singapore The Straits Times
- STExplains: Can customers say no to the use of their personal data for AI training?
STExplains: Can customers say no to the use of their personal data for AI training? The Straits Times
- AI-specific notifications mandatory for firms using personal data to train AI models: PDPC
AI-specific notifications mandatory for firms using personal data to train AI models: PDPC The Straits Times
- Vapi voice agent with AssemblyAI Universal-3.5 Pro Realtime
How to integrate Vapi with AssemblyAI’s real‑time model to build a voice agent.
- Entry-level chip hiring jumps 47% in AI boom
JobKorea said Monday that entry-level hiring in South Korea's semiconductor industry surged in the first half of the year, reflecting robust demand for talent as the AI-driven chip boom continues. According to the recruitment platform's analysis, job postings for entry-level semiconductor positions jumped 47 percent from a year earlier. Entry-level roles accounted for 12.2 percent of all semiconductor job postings, up 2.4 percentage points from a year earlier and the highest share in five years,
- Pipecat voice agent with AssemblyAI Universal-3.5 Pro Realtime
Creating a Pipecat voice agent using AssemblyAI’s real‑time speech model.
- Twilio phone agent with AssemblyAI Universal-3.5 Pro Realtime
Building a Twilio‑based phone agent powered by AssemblyAI’s real‑time speech model.
- Node.js voice agent with AssemblyAI Universal-3.5 Pro Realtime
Step‑by‑step tutorial for building a voice agent in Node.js with AssemblyAI’s real‑time model.
- Google plans new chip to run Gemini models more efficiently
Google plans new chip to run Gemini models more efficiently, the Information reports
- Google is building a chip with Gemini baked into the silicon
Most AI chips are general-purpose. You load a model onto them, and they run it. Google is reportedly trying something stranger: a chip that is the model, with Gemini’s blueprint etched into the hardware itself. The project, informally called “Frozen v2,” was reported by The Information and picked up by Reuters and Bloomberg Law. Alphabet […] This story continues at The Next Web
- Google's "Frozen v2" chip reportedly bakes Gemini's architecture directly into silicon for efficiency gains
Google is developing "Frozen v2," a server chip that bakes the Gemini architecture directly into hardware. According to internal sources, it could be 6 to 10 times more efficient than current TPUs. Scheduled for 2028, the chip would drastically cut Google's AI inference costs and could give the company a price advantage over OpenAI and Anthropic. The article Google's "Frozen v2" chip reportedly bakes Gemini's architecture directly into silicon for efficiency gains appeared first on The Decoder .
- Azure touts trio of new AI instances powered by AMD Helios racks
AMD says it will start ramping shipments from Q3
- Alphabet Stock Jumps After Information Report of New Google Chip
Alphabet Stock Jumps After Information Report of New Google Chip The Information
- Alphabet stock pops on report it's developing a more efficient AI chip
The new AI chip, called "Frozen v2," would embed parts of Gemini's architecture directly into the silicon, according to the report.
- Google Shares Gain on Report of Chip to Boost AI Efficiency
Shares of Google owner Alphabet Inc. gained on a report that the company is developing a server chip designed to optimize its Gemini artificial intelligence model.
- Google is working on a new AI chip designed to make Gemini more efficient
Alphabet, Google's parent company, is reportedly working on a new chip designed to make its Gemini models run much more efficiently.
- Head of U.S. federal AI testing institute resigned after just three months
The Commerce Department confirmed Chris Fall's departure and said a new director will be announced in the coming weeks
- Scoop: Trump AI security agency head resigns
Chris Fall, the director of the Center for AI Standards and Innovation, is resigning just three months after taking over the federal AI testing institute, the Commerce Department confirmed on Monday. Why it matters: The abrupt departure comes as the Trump administration grapples with how to deploy AI safely and the agency hashes out standards. Commerce spokesperson Benno Kass confirmed Fall's resignation to Axios. Context: Fall was appointed in April to lead Commerce's Center for AI Standards and Innovation after the administration reorganized the agency formerly known as the U.S. AI Safety Institute. CAISI is responsible for developing AI testing and evaluation capabilities and supporting standards for advanced AI systems. What they're saying: "Following Chris's departure, NIST Director Dr. Arvind Raman will continue to oversee CAISI and will serve as Acting CAISI Director," Commerce spokesperson Kristen Eichamer said in a statement to Axios. Raman, a former Purdue University engineering dean, was sworn in as the director of the National Institute of Standards and Technology on June 30. What we're watching: While Raman will serve as acting director, the office will remain without permanent leadership as the administration debates its next steps on AI standards and oversight. What's next: The Commerce Department expects to announce a new director in the coming weeks. A Commerce official said Fall's appointment was always intended to be temporary and that Raman has been reviewing candidates over the past several weeks. Editor's note: This story has been updated with a statement from the Department of Commerce.
- Top Trump AI safety director leaves job after 3 months
Top Trump AI safety director leaves job after 3 months Business Insider
- Trump administration's head of AI safety agency resigns after 3 months on job
Arvind Raman, the director of National Institute of Standards and Technology, will serve as acting director of CAISI, according to a spokesperson
- Head of US AI safety agency resigns
Head of US AI safety agency resigns Reuters
- Google plans new chip to run Gemini models more efficiently, the Information reports
Google plans new chip to run Gemini models more efficiently, the Information reports Reuters
- Trump’s latest AI czar has already resigned
The director role for the Center for AI Standards and Innovation (CAISI) has become a revolving door since David Sacks left his position as czar.
- Alphabet stock pops on report Google is building a Gemini-specific AI chip
The chip, internally dubbed "Frozen v2," could serve 6 to 10 times more tokens per unit of power than Google's latest TPUs
- Invven
Ai receptionist with booking GPS Planner
- Google plans new chip to run Gemini models more efficiently: Report
Google expects new chip, dubbed "Frozen v2," to help address an AI computing capacity crunch that has fueled internal tensions and prompted Google Cloud to decline deals with outside customers
- ChatGPT
Ask anything, create anything, and move from idea to action.
- Nib
Record, transcribe & turn meetings into AI notes
- AI’s Wildest Week: Models, And Lawsuits
OpenAI, Anthropic, xAI, Apple, and Google all made moves
- Loova Ads Studio
Create AI ads that convert, not just look good