AI News Archive: July 17, 2026 — Part 6
Sourced from 500+ daily AI sources, scored by relevance.
- When Flock Comes to Your Town: I Asked Experts What to Do About These AI Cameras
Flock is setting up surveillance cams and drones in cities nationwide. Citizens are fighting back. Here's everything you should know.
- Forget AI training data. This startup learned from slime mold
In a classic experiment more than a decade ago, researchers in Japan gave a slime mold—a single-celled organism known for forming efficient networks—a “map” of the cities around Tokyo. They represented each municipality with an oat flake, the organism’s favorite food, and watched as it created a network that looked eerily similar to the Japanese rail system. The takeaway was clear: Slime mold is surprisingly good at creating complex and efficient paths. A startup called Mireta Urban Dynamics is now using the same basic approach in software designed for urban planners. [Image: Mireta Urban Dynamics] “Very early on in my architecture degree, I was fascinated with strategies that nature and biology evolved for solving challenges,” says Mireta cofounder Raphael Kay. Humans have been designing transportation networks for thousands of years, but some organisms “have been solving analogous challenges for hundreds of millions, if not billions, of years,” he says. [Image: Mireta Urban Dynamics] Instead of training AI on the slime mold, the software copies the basic way the organism grows. “This is already a form of intelligence that has been evolved over a large number of evolutionary cycles,” Kay says. Then the software adds in multiple other layers. For example, a city planner looking at a subway network could add details about the city’s population distribution—so the tool prioritizes certain areas—and a flood map that helps it avoid other areas. (AI can help build these other features, though the core of the tool is based on biological intelligence, not AI.) [Image: Mireta Urban Dynamics] The tool can be used either to design transportation networks from scratch or to suggest small changes to existing networks. Mireta has begun working with design firms on a handful of projects, from a road network on a college campus to a new metro network. So far, those designs are still in the proposal stage with clients, though Kay expects them to move forward. [Image: Mireta Urban Dynamics] There’s already evidence that the tool works. In the pilot projects, the company says that the tool has created networks that are 20% to 30% more resilient for the same unit cost as alternatives. Resilience, in this case, means that if a disaster shuts down some part of a transportation network, people still have other ways to get around. That’s a challenge that slime molds know how to address. “There’s this beautiful naturally evolved trade-off between cost and resiliency that I think biology has had to solve because there’s a huge penalty for dying,” Kay says. “They’ve evolved strategies for strategic redundancy.” [Image: Mireta Urban Dynamics] As climate change increases disaster risk, from wildfires to flooding, planners are looking for ways to build more resilience into design. “It’s only recently that the winds have shifted, where municipalities are starting to really pay attention to and pay for long-term resiliency solutions,” Kay says. “What’s emerged is that there’s a lack of tooling and quantitative approaches to meet the demand of resiliency planning and network design. What’s really fascinating is that these organisms have been solving this challenge for so long.”
- Brilliant AI Solution to High Solar Soft Costs & Overpricing in the United States
For more than a decade, it’s been clear that Americans pay much more for rooftop solar power than Europeans, Australians, and others. We have much higher “soft costs,” including much higher “customer acquisition” costs. Part of it, as well, is the sales process and sales tactics leading to … not ... [continued] The post Brilliant AI Solution to High Solar Soft Costs & Overpricing in the United States appeared first on CleanTechnica .
- Announcing the Corrigibility Research Fund
TLDR: I'm managing a new fund, housed at Lightcone Infrastructure, that will award at least $200,000 in grants and prizes for corrigibility research in 2026 . Roughly half will go to traditional grants (first application deadline August 23rd ) and half for prizes recognizing excellent work done this year. If you have interest in working on corrigibility, now is a good time to start! Apply via email: grants@corrigibilityresearch.org Why this fund exists When I first dived into AI safety and alignment in 2009, the field was basically nonexistent. I've been relieved and gratified to see attention and funding grow, especially in the past few years. But even now, nearly all AI safety funding goes to evals, control, or interpretability. Work on alignment itself still remains deeply neglected, and it's only through alignment research that the core problems get solved. At this year's LessOnline I was talking about this dynamic with Peter McCluskey, particularly around our shared interest in corrigibility. In the wake of that conversation, Peter, being a long-time patron of alignment work, directed a portion of his philanthropy towards launching this fund, with the goal of increasing the amount of corrigibility research happening around the world. Lightcone Infrastructure agreed to house the fund and appointed me as its manager, due to my expertise on the subject. Why corrigibility Corrigibility is one of (if not the most) promising angle on creating superhuman artificial intelligences that reliably act in alignment with human values. Training for ethical behavior and direct alignment with humanity runs headlong into known challenges: prosaic methods can't reliably distinguish reward proxies from true goals, instrumental convergence means that even partly-aligned agents will become self-preserving and subversive, [1] and the philosophy of ethics remains woefully unsolved, such that we wouldn't even know what values to instill, even if we could reliably write the AI's values by hand. Corrigibility, by contrast, offers a solution: build an AI that aims to keep the human principal in the driver's seat, empowering them to make wise choices (perhaps aided by the AI's counsel), rather than relying on the AI's direct judgment. This runs the risk of concentrating power in the hands of humans who might misuse it, but human alignment is a less-fraught problem, and is amenable to known strategies, such as democratic oversight. Thanks to the nature of corrigibility, a purely-corrigible agent can be expected to avoid scheming and other instrumentally-convergent strategies. And while corrigibility itself does not solve the limitations of machine learning, it is a simpler target than all of morality, and there are reasons to hope that in practice, imperfectly-corrigible agents still cooperate with their principals to surface their flaws and assist in pushing towards even more corrigible assistants. This robustness gives hope in something closer to an iterative approach, where control and interpretability techniques come together to produce a realistic plan for scaling up to the level of human-intelligence and beyond. (For more of my thoughts, see CAST: Corrigibility As Singular Target ) I am not alone in placing a high level of emphasis on the need for corrigibility. Eliezer Yudkowsky and Paul Christiano have both written at length about how it is central to their best hopes for alignment, as well as many other brilliant alignment researchers. [2] The latest constitution from Anthropic mentions the term sixteen times, and has a dedicated section for it. OpenAI is similarly bullish on creating AIs that are tool-like and deferent. ( Arguably more so than Anthropic!) Despite this, the number of people working directly on corrigibility, such as on clarifying the concept, formalizing it, testing whether and how it can be trained into current systems, and mapping where it breaks, is vanishingly tiny. My hope is that this fund shifts that, both by directly paying for work and by broadly signaling that the work is valuable. What counts as corrigibility research For the purposes of this fund, anything that predictably advances humanity's understanding of the subject is fair game. This spans the full range from pure theory (e.g. formal models, impossibility results, decision-theoretic analysis) to pure empirical work (e.g. training experiments, evaluations of corrigible behavior in frontier models, surveys of how laypeople think about the topic). Distillation of existing work is also welcome. The fund will be prioritizing efforts that cut to the heart of the subject, but feel free to apply for funds even if your research is only tangentially related. The goal is to impact the AIs that actually get built. We're looking for work that is legible and relevant to the people making decisions about real systems. Theoretical work that's judged as too esoteric to be of interest to someone like Joe Carlsmith is unlikely to get funding. Work that's incompatible with mainline capability techniques (e.g. machine learning, transformers) is similarly unlikely to be greenlit by this fund. [3] We won't fund work that, in expectation, notably accelerates AI capabilities. The frontier labs are already doing more than enough to fund work that pushes us towards the brink. If you think your research accelerates things, but also makes progress towards corrigibility, feel free to reach out, but I am likely to point you elsewhere. Work that engages with corrigibility's risks and downsides is encouraged. Corrigibility has known risks and problems, and I want the field's understanding of these downsides to grow alongside work towards showing its promise. Work that presents corrigibility in an overly rosy "everything is safe/fine" way is less likely to get funding, as it might promote a false sense of security, and thereby push the world in a bad direction. [4] You do not need to agree with my particular framing of corrigibility (i.e. CAST) to get funded. Serious engagement with other framings — including arguments that those framings are better — is welcome. Grants and prizes The fund plans to disburse money this year through two general mechanisms: Prizes (>$100k). Retroactive awards for excellent corrigibility research done in 2026 : $40k awarded at the end of September At least $60k awarded in mid-December Prizes require no application. I'll be watching LessWrong, the Alignment Forum, arXiv, and elsewhere. Nevertheless, please send me pointers to corrigibility work (yours or others') that you think ought to be rewarded. Excellent work will be eligible to win prize money multiple times, including potentially in future years. [5] Grants (>$100k). Traditional, apply-in-advance funding for prospective work on corrigibility. To balance getting funds to people sooner and giving more time to prepare, the plan is for there to be two application rounds this year: Round 1: applications due August 23rd, 2026 Round 2: applications due October 31st, 2026 I encourage applicants to be ambitious and ask for however much would actually change their research trajectory towards corrigibility, but I expect typical grants to be around $5k–$35k, buying time for a focused project, a research sabbatical, compute for a mid-sized training run, etc. Grantees should use the funding to begin work this year, but research takes time and it's acceptable to not expect results until 2027. The hope is that work can get off the ground via a grant, and then supported more fully by retroactive prizes once it has been proven to be high-quality. Grants may be supplemental to other funding, such as salaries, other grants, and (of course) prizes. Why lean so hard on prizes? Prizes are results-oriented, rewarding work that actually happened and can be more clearly seen as high-quality. In some cases, prizes buy more research effort per dollar than traditional grants, encouraging a wide range of people to think about corrigibility and whether they have research ideas that might win. Prizes let researchers be rewarded without the overhead of a grant application. Prizes create opportunities to publicly honor good work, bringing attention to the best ideas and the people that developed them. How to apply To apply for a grant (or bring attention to work that might be prizeworthy), simply send an email to grants@corrigibilityresearch.org . The application process is deliberately lightweight and flexible. Tell me what you want to do, and what level(s) of funding you're hoping for. Detailed applications are more likely to get funding insofar as the detail helps demonstrate the worthiness of the work. If something is under-specified, I'm capable of asking follow-up questions. Once you know what you're hoping to do, if writing the application takes you more than a few hours, something has probably gone wrong. Miscellaneous fine print The Corrigibility Research Fund is a program of Lightcone Infrastructure Inc. All disbursements are grants made by Lightcone. My funding decisions are formally recommendations; Lightcone retains final approval on every grant and prize, and controls the fund's assets. I do not represent Lightcone. I serve as an unpaid volunteer. MIRI, my employer, is aware of and supports my role as fund manager, but is not otherwise involved. To avoid conflicts of interest, I won't be awarding prizes or grants to myself, my family, or anyone who is currently at MIRI, regardless of merit. All grants must serve the long-term future of humanity, writ large. No funds may be used for lobbying, political campaign activity, or private benefit. Hope My dream is that a year from now, as a result of this fund, there will be several people who think of corrigibility as their subfield , and will be able to proudly say that they were authors of prize-winning research that moved humanity closer to handling the question of how to make sure the transition to the age of thinking machines goes well. This research, ideally, then goes on to influence the researchers and engineers at frontier labs in years to come, helping them ensure the first artificial general intelligences are corrigible and safe. Let's get to work! ^ A perfectly aligned AGI might, for example, scheme against its creators and escape control so that it can save more lives and generally do more good in the world. ^ See Existing Writing on Corrigibility for some of the main commentary as of 2024. I also have a 2024 bibliography here . ^ If you have a corrigibility idea that depends on an alternative architecture or otherwise ML-incompatible strategy, it may still be of interest and worthy of funding from other sources. Feel free to email me at max@intelligence.org . ^ That being said, don't feel the need to distort your perspective towards doom when applying, or dress up your work in deliberately critical language. The most important criteria by far is the quality of object-level insight, not the tone. The warning about overly-rosy portrayals is more about setting a baseline for where I'm coming from. ^ The long-term financial existence of the fund is not guaranteed, but my intention with prizes like these is to retroactively reward people who direct their attention towards corrigibility. As such, if the fund continues into 2027 and beyond, I intend to allocate some prize money to work done in previous years. The Corrigibility Research Fund would especially love to encourage researcher-investor partnerships that use an impact-certificate-like model . Discuss
Score: 45🌐 MovesJul 17, 2026https://www.alignmentforum.org/posts/FBqe5dt8ZjaHN4Xj9/announcing-the-corrigibility-research-fund - Google’s AI just recreated the best goal ever by Pele that was never actually filmed
Google used live stadium capture, and combined it with archival material to bring Pele's legendary goal to life. The Veo video engine, Gemini Omni, and Nano Banana Pro AI tools tagged alongside in the journey to revive a lost marvel.
- AI Bubble Fears Are Starting to Spill Over
Investors are freaking out. The post AI Bubble Fears Are Starting to Spill Over appeared first on Futurism .
Score: 45🌐 MovesJul 17, 2026https://futurism.com/artificial-intelligence/ai-bubble-fears-tsmc-nvidia-earnings - ‘They’re going to kill me’: A Waymo rider was trapped inside as vandals smashed the robotaxi
‘They’re going to kill me’: A Waymo rider was trapped inside as vandals smashed the robotaxi San Francisco Chronicle
Score: 45🌐 MovesJul 17, 2026https://www.sfchronicle.com/sf/article/waymo-attack-rider-trapped-inside-22348656.php - Huge AI data center rejected in Florida. Why are they so controversial?
Huge AI data center rejected in Florida. Why are they so controversial? USA Today
- AI Was Supposed to Replace Court Reporters. The Data May Tell a Different Story.
AI Was Supposed to Replace Court Reporters. The Data May Tell a Different Story. USA Today
- Drones, AI and White Paint: Europe Races to Protect Infrastructure From Heat
As Europe’s railways buckle under record heat, roads melt and power grids strain, countries are turning to an array of fixes for aging infrastructure, from drones inspecting tracks and AI-powered sensors to a surprisingly simple tool: white paint. At Norway’s …
Score: 45🌐 MovesJul 17, 2026https://www.insurancejournal.com/news/international/2026/07/17/877881.htm - How Founders Can Rebuild Entry-Level Work for the AI Era
Companies still need to turn inexperienced hires into trusted decision-makers.
Score: 45🌐 MovesJul 17, 2026https://www.inc.com/haotianbai/how-founders-can-rebuild-entry-level-work-for-the-ai-era/91375076 - Tec-Do Launches Navos 2.0 at WAIC, Advancing the Next Era of Agentic Commerce
Tec-Do Launches Navos 2.0 at WAIC, Advancing the Next Era of Agentic Commerce The Straits Times
- Netflix dumps gRPC for SSE, eyes more AI
Any company that decides to build a graph database alternative on Cassandra probably has novel ideas worth listening to...
Score: 45🌐 MovesJul 17, 2026https://www.thestack.technology/netflix-dumps-grpc-for-sse-eyes-more-ai/ - Ex-MBDA engineer builds $1,100, 40g AI micro-drone using AMD phased-array sonar for YC-backed startup Tornyol to eradicate mosquitoes, currently restricted to 3-minute flights
Former MBDA engineer Alex Toussaint creates an AI mosquito-hunting drone using sonar, though battery limits restrict current operations.
- The Future of AI Infrastructure with CoreWeave
As AI applications become more complex, the infrastructure powering them needs to evolve. Corey Sanders, SVP of Product at CoreWeave, joins Chris to discuss why AI requires a fundamentally different approach than traditional cloud computing. They explore AI-native infrastructure, training and inference workloads, the rise of agentic development, optimizing GPU performance, AI research workflows, and why the future of software will be built around AI-first experiences rather than websites and apps. Featuring: Corey Sanders – LinkedIn Chris Benson – Website , LinkedIn , Bluesky , GitHub , X Links: CoreWeave Sponsors: Framer: The enterprise-grade website builder that lets your team ship faster. Get 30% off at framer.com/practicalai Upcoming Events: Register for upcoming webinars here ! Midwest AI Summit 2026
- tiket.com and Microsoft Bring Seamless Travel Services to Life with AI
The post tiket.com and Microsoft Bring Seamless Travel Services to Life with AI appeared first on Source .
- South Korea-US team unveils robotic technology that dresses the wearer
The technology developed by researchers at South Korea's KAIST and Stanford University uses soft and flexible "vines" powered by air pressure embedded in clothing. When pressurised, the vines glide the fabric up close to the wearer's body like an ivy plant climbing on a structure, even if the person does not remain standing still.
- AI teaches a bitter biology lesson
Human knowledge is getting in the way of scientific progress.
Score: 45🌐 MovesJul 17, 2026https://www.semafor.com/article/07/17/2026/ai-teaches-a-bitter-biology-lesson - Premio Expands Strategic Collaboration with Intel to Support Edge AI and Intelligent Automation
Premio Expands Strategic Collaboration with Intel to Support Edge AI and Intelligent Automation azcentral.com and The Arizona Republic
- Brex built its AI agent policy by watching what agents actually do, not by writing rules first
OpenClaw has become one of the most widely adopted agentic frameworks, but it has yet to prove itself at enterprise scale. Agents need real credentials — API keys, OAuth tokens, service accounts — to work effectively, and Brex found that traditional guardrails couldn't contain what those agents were doing with them. Brex set out to overcome these limitations by building an internal platform it calls CrabTrap. The open-source HTTP/HTTPS proxy intercepts all network traffic, examines policy rules, and uses a LLM-as-a-judge to decide whether agent requests should be approved or denied. “What we noticed was that the network layer was an untapped enforcement point,” Brex co-founder and CEO Pedro Franceschi told VentureBeat. “Every request an agent makes is an opportunity to intercept, reason about, and make a policy decision.” The takeaway Franceschi wants IT leaders to draw: agent governance should shift from SDK-level permissions and model guardrails toward a centralized network control plane that enforces and learns from real in-the-wild agent behavior. How Brex targeted the transport layer The “obvious fix” (at least initially) to the agent security gap was guardrails, and much of the early work has centered on scoped tools, per-action permissions, and human-in-the-loop approvals. But as agents evolve, each new capability means there’s another API to tune or surface to audit, Franceschi noted. “Any agentic system with multiple tools and access to the open internet creates an immediate tension for builders: The more capable you make an agent, the more dangerous it becomes, and the safer you make it, the less useful it is,” he said. Existing solutions to this tradeoff were “weak”: Fine-grained API tokens help at the margins but can still be misused and constrain functionality. Semantic guardrails (such as context, skills, or prompt steering) are easily bypassed by prompt injection, especially for agents connected to the internet. Agents can be “defanged” when given read-only access or limited toolsets, but then they can't do meaningful work, Franceschi said. On the other hand, granting broad write access and a large tool surface can result in hallucinations and real production consequences. Model context protocol (MCP) gateways enforce policy at the protocol layer — but only for traffic using MCP. Meanwhile, guardrails from LLM providers are tied to a single model and can be “opaque” to customize with enterprise-specific policies. And powerful tools like Nvidia OpenShell offer more of a “per-sandbox egress control.” “When we started, we hadn’t found a solution to deploying harnesses like OpenClaw safely,” Franceschi said. “Instead of waiting for the industry to catch up, we decided to own the problem and invent the necessary tools.” Notably, they needed a platform that sat between every agent and every network request, and could make “nuanced decisions about what to allow,” he said. This made the transport layer a core architectural component and natural starting point, he said. By operating at this layer, CrabTrap is framework-agnostic, language-agnostic, and API-agnostic. It doesn't require SDK wrappers or per-tool integration. Users set HTTP_PROXY and HTTPS_PROXY in the agent's environment, and every outbound request routes through the proxy before it reaches a destination. However, Franceschi emphasized, Brex didn't start at the transport layer because it thought it was the only answer; rather, they believe in “security by layers.” “The transport layer was simply an underinvested one, and we saw an opportunity to add meaningful enforcement there alongside everything else,” he said. The LLM-as-a-judge training loop CrabTrap combines deterministic static rules with an LLM-as-a-judge for requests that fall outside known patterns, Franceschi explained. The judge only “fires on the long tail of unfamiliar endpoints or unusual request shapes,” which for a mature agent is typically fewer than 3% of requests. The more pressing problem was how to know that a policy is the right one? With static rules, it's “relatively straightforward” to reason about accuracy. But with an LLM judge, the system is nondeterministic, and users need confidence that the policy approves the right requests and blocks the rest. “Our key insight was to bootstrap policy from observed behavior rather than write it from scratch,” Franceschi said. Beginning with real behavior and editing down based on real-world learnings turned out to be “dramatically more effective than starting from a blank page.” Brex’s team built a policy builder (itself an agentic loop) that runs underlying agents in shadow mode, analyzes historic network traffic, samples representative calls, and drafts a natural-language policy that matches what the agent actually does. From there, they built an eval system that tests policy changes before they go live. CrabTrap compares historical audit entries against a draft policy and reports the exact changes to be made. Users can slice results by method, URL, original decision, and agreement status. All of this runs with concurrent judge calls, so replaying thousands of requests “takes minutes, not hours,” Franceschi said. Brex also developed a live feedback loop: Full audit trails are stored in PostgreSQL and queryable through the admin API and dashboard. In cases where a resource is continuously denied, the system can notify a human or an agent to propose a policy update for review. “That closes the loop between observed denials and policy refinement,” Franceschi said. Core challenges and roadblocks Of course, the build wasn’t without its challenges. A big one was latency: “Putting an LLM between an agent and every outbound API request sounds like it would grind things to a halt,” he said. However, it didn’t turn out to be as big a problem as expected. This was for two reasons: The LLM judge only activates on a small fraction of requests (the aforementioned 3%). Agents quickly settle into predictable traffic patterns; once observed, high-volume patterns become static rules. Second, by using small, fast models like Claude Haiku meant that, even when the judge did fire, added latency was “negligible.” This can be further reduced with local models and prompt caching, Franceschi said. The harder and less obvious challenge was prompt injection, he said. The judge receives the full HTTP request and all content is user-controlled, so potentially, a crafted URL, header, or request body could manipulate the judge's decision. Brex addressed this by structuring the request as a JSON object before sending it to the model, so all user-controlled content is “escaped rather than interpolated as raw text,” Franceschi said. Results, and where CrabTrap might evolve Brex tracks a few factors to measure CrabTrap’s internal impact: Engagement with agents, network traffic patterns, and net promoter scores (NPS). The most meaningful result of CrabTrap has been “organizational confidence,” Franceschi said. Previously, the team had “real hesitation” when it came to deploying autonomous agents broadly across business operations, because the existing guardrail options didn't provide enough assurance. “CrabTrap changed that calculus,” Franceschi said. They now have an enforcement layer they trust, increasing confidence around expanding agent deployment into more parts of the business and delegating more agent configuration and management to users. Franceschi described the policies derived from traffic as “surprisingly strong.” The team expected the policy builder to produce a “rough starting point” requiring heavy manual editing. In practice, though, pointing the platform at a few days of real traffic produced policies that matched human judgment on the “vast majority of held-out requests.” Additionally, CrabTrap revealed how much noise agents generate. “The audit trail made this visible for the first time,” Franceschi said. They used denial logs and traffic analysis not only to tune policies, but to tighten agents themselves, remove tools, and cut out entire categories of requests that were wasting both time and tokens. “The proxy became a discovery tool, not just an enforcement one,” he said. Areas for growth (and input from the open-source community) Brex anticipates CrabTrap to continue to evolve, particularly as they have released it as open-source. “We hope the community helps shape it,” Franceschi said. Areas of improvement include deeper authentication functionality such as single-sign on (SSO), fine-grained role-based access control (RBAC); escalation workflows that allow agents to request additional permissions; and policy recommendations based on denial patterns. Programmatic configuration, or developing API endpoints for “creating, forking, and applying” policies to agents, could allow the whole policy lifecycle to be automated rather than managed manually, Franceschi said. As for escalation, if an agent is continuously denied a given resource or endpoint, it should be able to route requests to humans or other AI agents for review and back that up with a rationale for why it needs access. “That turns CrabTrap from a hard enforcement boundary into something more like a managed permission system,” Franceschi said. Additionally, the policy was built to bootstrap from network traffic, but there is opportunity to incorporate additional signals around agent traces and resource-calling, as well as broader context on what agents are ultimately trying to accomplish. This can help produce more accurate and nuanced policies. Finally, there's an “open philosophical question” about the right posture for CrabTrap: Should it be a fully transparent layer that the agent itself is unaware of, or should it operate more like a “well-intentioned manager”? (that is, the agent knows about the layer and can interact with it). The open-source community can help shape these developments, and CrabTrap will only get better with more users, Franceschi said. Brex’s agents speak to a specific set of APIs; teams using CrabTrap with different agents, services, and policy requirements will surface “edge cases and patterns we can't hit alone.” “We have ambitious plans for where it could go, and we’d rather build in the open,” Franceschi said. What other builders can learn from CrabTrap The response has been stronger than expected. CrabTrap has more than 700 stars on GitHub . Franceschi said Brex has also heard from OpenAI, Y Combinator CEO Garry Tan, and programmer Pete Steinberger, all expressing interest in deploying similar internal infrastructure. The broader lesson: “Don't let infrastructure gaps become excuses to wait," Franceschi advised. There are “real blockers” for every enterprise looking to seriously deploy AI agents, including security concerns, lack of tooling, or unclear guardrails. “It's tempting to sit on your hands until the industry catches up,” he said. “The lesson from CrabTrap is that you can own those problems directly.”
- X cracks down on content theft, AI now spots copied posts three times faster
X cracks down on content theft, AI now spots copied posts three times faster
- Myntra Scales AI Across Customer Discovery, Seller Onboarding and Product Development
Myntra today outlined the next phase of its technology buildout, anchored by AI capabilities spanning customer discovery, seller onboarding, catalog infrastructure, and how its engineering teams build and ship products. The rollout is built around three core pillars — customer experience, seller enablement, and operational efficiency. Across each pillar, AI is doing specific, measurable work. […] The post Myntra Scales AI Across Customer Discovery, Seller Onboarding and Product Development appeared first on CXOToday.com .
- Pilots and airlines are using AI tools to better time the seatbelt sign
Pilots and airlines are using AI tools to better time the seatbelt sign Business Insider
Score: 43🌐 MovesJul 17, 2026https://www.businessinsider.com/pilots-airlines-using-ai-tools-better-time-seatbelt-sign-on-2026-7 - 5 steps to secure your infrastructure in the frontier model era
The industry conversation around AI infrastructure has narrowed to a single dimension: scale. The focus is on GPUs, power, cooling and the massive physical footprint required to train and run AI agents and models. At the same time, organizations are adjusting to the speed and scale with which AI is identifying vulnerabilities — which is much faster than remediation can be started. However, almost no one is talking about the infrastructure layer that actually determines whether AI workloads remain secure, resilient and compliant. This is the layer that runs the world’s most sensitive, regulated, high‑value workloads. Thankfully, it already has the guardrails needed for an era where vulnerabilities are discovered faster than ever. But are they being set correctly? With more than one billion AI agents expected by 2029 , organizations need a plan for their infrastructure layer to withstand threats from new frontier models, maintain uptime and protect data sovereignty. As they scale AI deployments, enterprises must secure the infrastructure AI depends on. These five steps outline what organizations can do now to strengthen their infrastructure posture using proven, enterprise‑grade practices for current and future threats. Step 1: Build on infrastructure engineered for security and resilience Infrastructure must be secure by design, not secured after deployment. The systems that have historically supported the world’s most critical workloads — from global payments to national‑scale operations — were built with this principle at their core. If you’ve already invested in systems designed for mission-critical workloads, you’ve checked this first box. Enterprise‑grade systems have been engineered with multilayered security controls, pervasive encryption, confidential computing and hardware‑level protections that make exploitation dramatically harder. A frontier model in the hands of a bad actor can chain weaknesses faster than humans can patch them — unless the underlying infrastructure is built to absorb and deflect that pressure. When I meet with clients, I often tell them what our own security teams operate under: we assume vulnerabilities will continue to be discovered and we design for that reality. That mindset is what separates infrastructure that survives frontier‑model pressure from infrastructure that collapses under it. These systems continue to evolve with predictive failure analysis and accelerated recovery, allowing systems to continue operating even during investigation and remediation. Step 2: Treat uptime and resilience as a security requirement If your infrastructure fails, your workloads will too. These systems depend on uninterrupted access to data and compute, and even seconds of downtime can compound operational and security risk. Enterprise‑grade platforms deliver near‑continuous availability through redundant hardware paths and intelligent system recovery. The easiest fix? Ample resources and an up-to-date infrastructure foundation. Too often, a security problem is really an availability problem that turned into a security problem. When systems fall behind on maintenance, capacity or recovery readiness, they create the exact openings a frontier model can exploit. A delayed maintenance cycle or a recovery process that takes too long becomes the opening a frontier model can exploit. Resilience is not just about uptime. It is a security control. And this will not be the last time a frontier model tests the limits of that resilience. Data resilience is equally critical. Cyber‑resilient storage systems with immutable backups and rapid recovery capabilities ensure that critical data remains protected and available even after a cyber incident or disaster. Step 3: Operate for continuous discovery, not periodic defense The idea that you can prevent every vulnerability is outdated. The more realistic model is continuous discovery — finding, prioritizing and addressing issues faster than they can be exploited. Organizations must operate as if vulnerabilities will be found faster than ever. Instead of relying on static defenses, they should emphasize layered controls, rapid triage, continuous delivery of fixes and coordinated disclosure. Frontier models in the hands of bad actors can amplify security challenges by connecting vulnerabilities. They can chain misconfigurations, outdated components and privilege gaps into a viable attack route in minutes. And the more outdated or inconsistent an environment is, the easier that chaining becomes. Modern operational‑intelligence tooling helps them surface that risk, prioritize what matters and act before an attacker can exploit the gaps. These platforms help organizations understand where they are exposed, identify which maintenance issues carry the highest operational and security risk, and reduce the blind spots that frontier‑model attackers are increasingly adept at exploiting. It’s critical to assess how you manage your vulnerabilities. Internal processes should address severe vulnerabilities within hours, regardless of whether they are discovered by humans, traditional tooling or AI‑driven techniques. As AI accelerates vulnerability chaining, this posture maintains operational integrity and reduces exposure. Step 4: Use AI to defend AI Leading organizations are integrating AI‑driven threat detection directly into their infrastructure. On operating systems like z/OS, AI‑based analytics can identify anomalous and potentially malicious data access, reducing investigation time and limiting impact. Beyond detection, autonomous security models are emerging that continuously govern risk, investigate threats and enforce resilience across identities, data, applications, cloud and networks. Across the industry, we’re seeing the rise of autonomous security frameworks that use AI to assess posture, detect threats and harden controls without waiting for human intervention. Combined with modern AI‑accelerated processors, these capabilities allow threats to be analyzed and mitigated directly within the infrastructure itself. Step 5: Join a broader ecosystem fighting frontier model threats No organization can face frontier model threats alone. These risks require coordinated industry action. Frontier models give both good and bad actors the ability to analyze codebases, chain vulnerabilities and probe infrastructure at a scale that no single enterprise can counter on its own. Across the industry, coalitions are emerging to assess and remediate vulnerabilities discovered by frontier-class models and to help enterprises build AI resilience. Initiatives like Project Glasswing, Project QuiltWorks and the Frontier AI Alliance are examples of how providers, consultancies and security firms are beginning to coordinate their response to AI-accelerated threats. Organizations can also benefit from independent assessments that evaluate readiness for agentic-enabled threats and identify gaps across their infrastructure. These assessments help teams understand where they are exposed, how frontier models might chain those exposures together, and what actions will reduce the likelihood of a high-impact event. Participating in these programs is one of the most concrete steps enterprises can take today to strengthen their AI infrastructure posture. Your AI security depends on the infrastructure you choose AI is accelerating both innovation and risk. The organizations that succeed will be those that build on resilient, secure infrastructure, prioritize uptime as a security control, operate with continuous discovery, use AI to defend AI and participate in the global response to frontier‑model threats. In the end, your ability to scale AI safely comes down to the infrastructure you trust to run it. This article is published as part of the Foundry Expert Contributor Network. Want to join?
Score: 43🌐 MovesJul 17, 2026https://www.cio.com/article/4197934/5-steps-to-secure-your-infrastructure-in-the-frontier-model-era.html - Strategic Simulation Is AI’s Next Frontier for Enterprise Decision-Making and Leadership
Strategic Simulation Is AI’s Next Frontier for Enterprise Decision-Making and Leadership uk.entrepreneur.com
- Rothman: Mass. shouldn’t let AI firms grade themselves
Rothman: Mass. shouldn’t let AI firms grade themselves Boston Herald
Score: 42🌐 MovesJul 17, 2026https://www.bostonherald.com/2026/07/17/rothman-mass-shouldnt-let-ai-firms-grade-themselves/ - Mira Murati’s 975B Inkling Doesn’t Beat GPT or Claude. That’s the Point.
Thinking Machines shipped it under the Apache 2.0 licence, the most permissive licence there is. It tops no leaderboard, because it was… Continue reading on Towards AI »
- How one bank is managing the risk that AI spending dries up
Amid warnings that a future slowdown in AI capital spending could pose a systemic threat, executives at Regions Financial said they're preparing the same way they would for any other credit concentration.
Score: 42🌐 MovesJul 17, 2026https://www.americanbanker.com/news/how-one-bank-is-managing-the-risk-that-ai-spending-dries-up - AI Security Is Never Finished: Building the Continuous Red Teaming Loop
Security programs are built around moments of closure: the finding is closed, the control passed, and the release can move forward. AI security refuses to cooperate with that model. A passing test is useful evidence, but only for a specific system, configuration, and moment. Production AI keeps moving. Models change behavior, prompts are revised, retrieval […] The post AI Security Is Never Finished: Building the Continuous Red Teaming Loop appeared first on CXOToday.com .
- Analyst Take: AI (Including AI Agent) Workloads Require an AI-Native Infrastructure
Analyst Take: AI (Including AI Agent) Workloads Require an AI-Native Infrastructure Gartner
- Why My LLM Guardrail Flagged the Right Answers (And Why I Refused to Fix It)
Benchmarking a numerical hallucination checker against a frontier API and a local 8B model taught me a harsh lesson about system design. Introduction Ask a Large Language Model (LLM) to write an executive strategy document based on your machine learning pipeline’s outputs, and sooner or later, it will invent a number. It will confidently claim a “47% lift” when your optimizer actually said 12%. In a boardroom setting, this single hallucinated metric is enough to burn the credibility of the entire dashboard. Due to this, most data teams either force a human into the loop for every generated figure, or they simply refuse to ship the LLM layer at all. I decided to take the third route while building GTM Wargame (an open source go-to-market strategy simulator). I wanted to treat “ don’t hallucinate ” as a systemic architectural property, rather than relying on prompt engineering tricks. To do this, every number the agents were allowed to use - SHAP attributions, optimizer outputs, market context was stored in a strict Ground Truth Pool. Before any text reaches the UI, a ConsistencyChecker cross-references every number the agents write against that pool. Building this guardrail was the easy part. The harder question was: How often does it actually fire in a real world scenario, and when it does, is it right? To find out, I ran 60 paired simulations pitting a frontier API model against a local 8B model, and I hand audited every flagged number against a deterministically rebuilt ground truth. The first finding was expected: the local model fabricated numbers at a materially higher rate (7.2%) . The second finding is what this article is really about: every single flag the guardrail raised against the frontier model was a false positive. The model wasn’t hallucinating. It was doing correct arithmetic that my checker simply had no way to understand. My first instinct as an engineer was to “fix” the checker. By the end of the audit, I realized those false positives are the guardrail working as designed, and expanding its matching logic to eliminate them would make the whole system less trustworthy. The guardrail pattern The guardrail’s logic is deliberately boring, which is exactly why it works: Build the Ground Truth Pool : Every number the agents might reference (SHAP values, market signals, optimizer results) gets flattened into one unified pool of floats. Let the agents write freely: The Analyst-Strategist-Manager chain produces unconstrained prose using chain-of-thought reasoning. There is no rigid JSON schema they have to fight against. Parse the output after the fact: A regex pass extracts every number the LLM generated and checks it against the pool. It allows for rounding, percentage-vs-fraction scaling (“12” for a stored 0.12), and k-notation (“176k” for a stored 176,432.11). Surface the verdict: An unmatched number is flagged as a hallucination in the UI. The user sees the warning as the system doesn’t quietly try to “fix” the figure behind the scenes. The GTM Wargame pipeline: Every layer upstream of the agents exists to populate the Ground Truth Pool that the guardrail checks against. (Image by Author) Before testing it on real models, I evaluated the checker itself as a classifier using 75 labeled cases built from authentic pipeline artifacts. The cases were split into clean cases, injected cases (seeded with unmatchable numbers), and trap cases (filled with naive-checker bait formatting like $1,234, negative signs, and k-notation of real pool values) . The checker scored a perfect 1.000 precision and 1.000 recall, with zero false positives across the 258 numbers it ran on. result = ConsistencyChecker.validate_response( text=manager_summary, shap_info=shap_values, market_context=market_context, opt_results=optimizer_output, ) # Returns: {"is_valid": False, "hallucinated_values": [47.0], "error_msg": "..."} However, a perfect score deserves skepticism. The labels for this evaluation were defined by the checker’s own matching rules, so by construction, this eval couldn’t disagree with itself. It validated everything layered around the rule - number extraction, currency and comma cleaning, sign handling, pool flattening. This however says nothing about whether the matching rule itself was the right architectural contract. That question needs real models. The Experiment: Two Models, 30 Paired Seeds (60 Runs) I ran the full pipeline (XGBoost forecasting, SHAP attributions, SciPy budget optimization, followed by the agent boardroom) across 30 identical seeds. The pipeline runs on a synthetic sales dataset generated by the repo itself, so no proprietary data is involved and every result can be rebuilt from a seed. Both models faced the exact same scenarios: Gemini (gemini-3-flash-preview, the repo’s pinned default) over the API, and Llama 3.1 8B Instruct (Q4_K_M) running fully locally through llama.cpp. These runs were recorded in July 2026 − preview models change over time, so the exact version and date matter for anyone comparing against these numbers. The checker scored the analyst and strategist nodes of every run. Before looking at the outputs, I established a strict methodological rule: A flag is not a fabrication. A flag means a number failed pool matching. That is an upper bound on hallucination, not a measurement of it. Reporting raw flag rates as “fabrication rates” is exactly the kind of unverified claim the guardrail exists to prevent. Here are the raw flag counts: Source: Image by the author. Two things immediately stood out. First, the frontier model engaged with the quantitative data far more densely, writing nearly four times as many numbers per run yet triggered fewer absolute flags. Second, 13 total flags across all runs is a small enough population to audit exhaustively by hand. Auditing the Flags: Mechanical Evidence & Judgment For every flagged run, I used an offline audit script with zero LLMs in the loop to deterministically rebuild that run’s exact Ground Truth Pool from its seed. It revalidates the stored agent transcripts, and asserts the flags reproduced exactly, ensuring I was examining exactly what the checker saw at runtime. Then, I ran two mechanical passes per flag: Notation Pass: Is the value a valid x1,000 or x1,000,000 scaling of a pool value? Derivation Pass : Does the value result from a pairwise combination of pool values (difference, sum, ratio, percentage gap, per week average, etc.)? The derivation pass comes with a built-in trap. With roughly 50 values in a pool, the number of pairwise combinations runs into the thousands. By pure chance, some derivation will land near almost any small number. Therefore, a mechanical hit isn’t enough. The final classification required semantic support from the text: the model had to explicitly name the operands it was combining. Any ambiguous flags defaulted to FABRICATED . Here are the final verdicts: Source: Image by the author. Classification of the 13 raw flags against the deterministically rebuilt ground truth pool. (Image by Author) Every single one of the local model’s 10 flags survived the audit as a genuine hallucination. Conversely, not a single one of the frontier model’s flags was an actual fabrication. The False Positives: Correct, But Derived All three of Gemini’s flags followed the exact same pattern. Let’s look at the flag from Seed 1015. The ground truth pool contained an optimized price of 535.68 and a market leader price of 715.18. The LLM strategist wrote: “ …our reliance on a 25 percent price gap relative to the leader (715.18) makes us highly vulnerable to any aggressive downward moves from the competition. ” Check the arithmetic: (715.18 − 535.68) / 715.18 = 25.1%. Both operands are real pool values. Both are named in the surrounding text. The model computed a mathematically correct, highly relevant business metric. And the checker flagged it, simply because the number “25” wasn’t explicitly stored in the original pool. Seeds 1019 and 1023 produced the same pattern with different numbers, each verifiably correct, each flagged. I call this category correct-but-derived . It’s a failure mode that standard guardrail evaluations miss entirely because the eval’s labels are defined by the checker’s own matching rules. It took real model behavior and a manual audit to expose the flaw. Why I Refuse to “Fix” It The obvious engineering patch writes itself: teach the checker arithmetic. Extend matching logic to cover pairwise derivations over ground truth pool values and all three of Gemini’s false positives would vanish. I actually built that logic (it’s the derivation pass used in the audit script), but running it offline convinced me to never put it into a live guardrail. Here is why: The collision problem: Thousands of pairwise combinations over 50 pool values create chance hits everywhere. The audit handles this by requiring semantic support from the sentence, which works when a human reads thirteen flags with full context. To prevent the checker from accidentally validating a hallucination at scale just because it mathematically collided with a random derivation, I would need to put an LLM inside the guardrail to judge the semantic context. At that point, the component whose only job is to catch LLM errors becomes dependent on an LLM being right. The asymmetry of the two failure modes: A false positive is cheap and transparent; a user sees a warning on a correct number, verifies it, and moves on. The system’s skepticism is visible. A false negative, however, is expensive and invisible: a hallucinated figure sails into an executive playbook wearing a “System Verified” badge. The entire value of a guardrail is that a green light signifies a strict, undeniable truth: this number exists in the raw data. Every time the guardrail matching contract is expanded, the value of that green light is silently diluted. So, the contract stays narrow. The checker verifies existence, not derivability. Deriving new true values is legitimate LLM behavior that falls outside the contract. It gets flagged, and the human user resolves it. The Grounding Tax Now, let’s address the local model, because this is where the guardrail proved it was heavily load-bearing. After the audit, the true fabrication rates separated cleanly: Source: Image by the author. Post-audit fabrication rates with 95% Confidence Intervals. (Image by Author) The character of the local model’s hallucinations is highly instructive. It didn’t just get numbers slightly wrong; it invented entire quantitative structures that the pipeline never produced. It hallucinated an entire ROI table − $100k spend, $500k revenue” − in a pipeline that computes neither revenue nor ROI. My personal favorite was from Seed 1017: “ Reduce budget by 20% to $X. ” The model fabricated a metric and left the template variable sitting right next to it in the exact same sentence unfilled. In my last article, I wrote about the “ Abstraction Tax ”, the throughput you sacrifice when using convenient local inference wrappers on constrained hardware. This is the other side of the coin: The Grounding Tax . A quantized 8B running on your laptop buys you complete data privacy and zero API costs. But in a highly analytical workload, it pays for that privacy with a 7.2% hallucination rate on executive-facing outputs. This in part is mitigated by using bigger models on better hardware. The guardrail firing on a live local-model run, catching an invented ROI fabrication. (Image by Author) This presents an operational paradox. The deployments that force engineers onto local models − regulated data, air-gapped systems, strict client confidentiality − are the exact environments where output verification is the hardest to outsource, and where a hallucinated number does the most damage. If your privacy requirements push you toward local inference, a grounding check isn’t just a nice-to-have feature − it is the only thing that makes the deployment defensible. Two caveats. First, this guardrail benchmarks numerical grounding against a known pool, not semantic correctness. Second, K=30 runs on a single local model quantization is enough to establish confidence intervals, but it is not a definitive industry leaderboard. The takeaway isn’t “Gemini beats Llama”; the takeaway is that fabrication rates vary wildly across deployment architectures, and you need to measure yours. Conclusion I set out to measure how often LLMs invent numbers when summarizing ML outputs, and I walked away with a system design principle I didn’t expect: the audit matters more than the guardrail . Raw metrics told a highly misleading story (0.6% vs 7.2% flag rates). Only a manual classification audit revealed the truth: a 0% vs 7.2% actual fabrication rate, where every single flag against the frontier model was a mathematically correct derivation that the checker was right to be suspicious of, but wrong to condemn. If you are building numerical guardrails for LLMs, here are the rules I’m carrying forward: A flag is an upper bound, not a claim: Audit your flags before reporting fabrication rates, and resolve ambiguities against the model. Keep the checker’s contract narrow: Verify existence, not derivability. A strict, trustworthy green light is worth the cost of a few honest warnings. Evaluate against real model behavior: The most interesting failure class in this system (the correct-but-derived false positive) was entirely invisible to my standard, perfect-scoring evaluation script. All code, raw data, and empirical setups used to generate these profiles are fully reproducible. The checker evaluation can be rebuilt offline without API keys, and every flagged transcript is available in the GTM Wargame repository . Why My LLM Guardrail Flagged the Right Answers (And Why I Refused to Fix It) was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.
- Vertu wants executives to pay $6,880 for an AI agent — here’s how it actually performs
From AI workflows to battery life and security, here's what it's really like to live with Vertu's luxury foldable every day.
- Claude Code’s creator offers a better way to measure AI success than token burn
Claude Code’s creator offers a better way to measure AI success than token burn Business Insider
Score: 41🌐 MovesJul 17, 2026https://www.businessinsider.com/claude-code-boris-cherny-better-way-measure-ai-success-dashboards-2026-7 - AI and the crisis of recognition: Do we still see the human behind the words?
I couldn’t understand why my student had ignored almost all of my feedback. I had carefully reviewed his capstone presentation, working through each slide, thinking about the structure, the technical flow and how he could communicate his ideas more effectively. Like many educators adapting to the AI era, I also used AI to help organise […] The post AI and the crisis of recognition: Do we still see the human behind the words? appeared first on e27 .
Score: 40🌐 MovesJul 17, 2026https://e27.co/ai-and-the-crisis-of-recognition-do-we-still-see-the-human-behind-the-words-20260714/ - AI didn’t replace our Security Team, it multiplied it
Webflow's security engineers built AI into triage and post-incident work. One change alone saved 504 hours in a single quarter.
Score: 40🌐 MovesJul 17, 2026https://webflowmarketingmain.com/blog/ai-didnt-replace-our-security-team - Engadget Podcast: Is Siri AI actually useful in iOS 27 and macOS Golden Gate?
Also, we try to make sense of OpenAI's rumored AI smart speaker.
Score: 40🌐 MovesJul 17, 2026https://www.engadget.com/2217377/engadget-podcast-is-siri-ai-actually-useful-in-ios-27-and-macos-golden-gate/ - How to catch voice agent regressions before your users do
Learn how to detect voice agent regressions early to improve user experience.
Score: 40🌐 MovesJul 17, 2026https://assemblyai.com/blog/how-to-catch-voice-agent-regressions-before-your-users-do - China’s AI Summer Camps Tap Into Parents’ Worries
As parents push to equip their children with AI skills, companies seeking to exploit their anxieties around the new technology are exaggerating claims about what they can teach their children.
Score: 40🌐 MovesJul 17, 2026https://www.sixthtone.com/news/1018774/China’s AI Summer Camps Tap Into Parents’ Worries - The build vs. buy dilemma at the heart of enterprise AI
For three decades, enterprise software has been a buy-it decision. Packaged software from SAP, Oracle and Salesforce covered roughly 80% of requirements at a fraction of the cost of building. The economics were obvious, and for traditional applications, they still are. AI is introducing a wrinkle that is forcing even the most committed enterprise software customers to rethink their options. AI is a layer that sits across your data, your processes, and your decisions. Where that layer runs and who controls it is an architecture question, and most of the enterprise community is still treating it as a procurement one. The appeal of vendor-embedded AI is clear: automated operational decisions, smarter supplier and merchandising choices, and friction-free workflows built into the systems enterprises already rely on. The catch is that these capabilities almost universally depend on your data living in the vendor’s cloud environment. For most large enterprises, it sits on-premises, in hyperscale cloud infrastructure they manage themselves, or in private data centers. That gap between where your data is and where your vendor’s AI assumes it should be creates a fundamental strategic fork in the road. Build vs. buy is a category error The framing I keep hearing is “build vs. buy your AI strategy.” It implies that some organizations are out there training foundation models from scratch. Nobody serious is doing that. The real choice sits across three distinct approaches, and conflating them leads to poor decisions: Buy embedded. Use the AI capabilities your vendor ships natively inside their platform: the assistant baked into your ERP, your CRM, your HCM suite. Lowest integration cost, fastest time to value, tightest fit with the application data. Buy platform. Adopt the vendor’s AI infrastructure layer and build your own assistants and agents on top of it. More flexible, but you remain inside the vendor’s architectural boundary and subject to their governance model. Compose. Connect a third-party model (Claude, GPT, Gemini, an open-weight model running in your own environment) directly to your existing landscape. Maximum control, maximum integration burden, and full responsibility for what comes out the other end. These are not equivalent options at different price points. They make different assumptions about where your data lives, who governs the AI, and how much architectural change you’ll absorb to get there. Vendor pitches sometimes blur the distinction on purpose. Enterprise leaders can’t afford to. The vendor AI stack has an assumption baked in Every embedded AI capability ships with an unstated architectural prerequisite: your data must be where the AI can see it, in the shape it expects, under the governance the vendor enforces. For organizations with clean, modern cloud estates, that is often a reasonable trade. For the long tail of large enterprises running heavily customized environments on private or hybrid infrastructure, that trade becomes a precondition, one you must meet before the AI conversation can even begin. Whether meeting it makes sense depends on your starting point, your sector’s regulatory posture, and your appetite for migration risk. None of those are uniform across organizations. That’s the part that gets glossed over in vendor keynotes. The AI demo on stage assumes a destination architecture the audience hasn’t necessarily reached yet. Large enterprise customers are carrying an unusually heavy technology burden right now. Many are simultaneously managing platform modernization programs that have been building for over a decade, alongside pressure to migrate to vendor-managed cloud infrastructure. Sitting above both is a boardroom-level directive to demonstrate meaningful AI progress fast. The vendor path to AI and the boardroom path to AI can diverge sharply, and enterprises need to make selective, strategic decisions about where to adopt AI first to maximize value and minimize risk. Sovereignty isn’t a slogan, it’s an architecture constraint The conversation about sovereignty has been hijacked by both sides. One camp treats every SaaS adoption as a sovereignty violation. The other dismisses every sovereignty concern as Luddite resistance. Neither is useful. What’s happening in real customer conversations – particularly in DACH, public sector, and financial services – is more specific. Organizations are drawing a distinction between running their applications in a vendor’s cloud (which is broadly fine, well understood, decades of precedent) and enriching their data and processes inside a vendor’s AI model (which has less precedent, is harder to reverse, and carries material implications for competitive position). Enriching your data inside a vendor’s AI model is the genuinely new question, and organizations that conflate it with their existing cloud posture tend to defend the wrong perimeter. Despite spending around $100 million annually with Amazon, Disney built its own internal AI system to house its corporate intelligence rather than rely on a hyperscaler’s AI offering. The decision came down to control. When your data represents decades of creative and commercial IP, you think carefully about where it lives and who can learn from it. Disney has become more open to SaaS over time. The AI sovereignty question is a separate debate from the SaaS debate and conflating the two leads organizations to the wrong conclusions. At the other end of the spectrum, enterprises in heavily regulated environments treat data sovereignty as an absolute non-negotiable. Any AI model must run within their controlled environment, especially where sensitive data cannot touch the public internet. GDPR obligations reinforce this instinct across the European market, requiring organizations to maintain clear accountability for how personal data is processed inside AI systems, including vendor-managed ones. AI-enriched data, meaning models that have learned the shape of your business processes, your supplier negotiations, your customer behavior, carries a different half-life and a different strategic value than the operational data underneath it. That deserves its own architectural decision, separate from your broader cloud strategy. What this means in practice Most large enterprise estates will end up with a mix of all three approaches, and where you draw the lines matters more than your overall posture. Embedded AI capabilities are the right answer for in-application productivity: the assistant inside your ERP workflows, the agent inside your procurement or HR suite. That is where vendor embedding genuinely shines, and attempting to compose your own equivalent is typically a poor use of engineering resources. Compose belongs elsewhere: in cross-application orchestration, in custom assistants over operational and observability data, and in agents that need to reach across multiple vendor systems and infrastructure layers in ways no single vendor stack will never natively support. Research from McKinsey suggests the most significant near-term productivity gains from enterprise AI will come precisely from these cross-system workflows, rather than from within individual applications. The most interesting enterprise AI work over the next eighteen months lives here, and it doesn’t require waiting for a migration to complete first. That compose path isn’t free, and it’s important to be honest about the costs. Governance, audit trails, and accountability for hallucinated outputs become your problem, not the vendor’s. Prompt drift and evaluation discipline are real engineering costs that never appear in the proof-of-concept. Those costs scale with the complexity of your landscape and the number of systems your agents touch. Budget for them before deployment, not after your first production incident. None of that is a reason to avoid the path. It’s a reason to staff for it, honestly. The real question The build-vs-buy frame survives because it gives executives a binary choice along a familiar axis. AI sits somewhere else entirely. The question worth putting on the table at your next architecture review is simpler: Which decisions do we want our vendors’ AI to make, and which do we want to keep on our side of the boundary? Answer that, and the right build/buy/compose mix flows from it. Skip it, and you will end up with the architecture your vendors prefer – which may or may not be the one your business needs. This article is published as part of the Foundry Expert Contributor Network. Want to join?
Score: 40🌐 MovesJul 17, 2026https://www.cio.com/article/4197957/the-build-vs-buy-dilemma-at-the-heart-of-enterprise-ai.html - Reasons to believe current AI models are conscious
There are a number of reasons to believe current AI models are conscious. I mean “conscious” is the sense of “is there something it is like to be an AI model?” and “does the AI model have phenomenal experience?”. As to what “AI models” refers to, the short answer is “y’know, like instances of Claude Opus 4.8 or GPT-4o”. [1] By “current AI”, I mean big post-trained LLMs. This piece was originally written as a document for myself and my friends. Imaginary interlocutors would ask me if I thought AI was conscious, and I’d say, “Probably, although ‘mu’ might be the better answer because I think we’re moving into territory where we lack the proper ontology and we don’t have the right concepts. [2] A more measured answer would be: 'current AI models/instances probably have the thing that ‘consciousness’ is pointing at in the important sense. Note though that their qualia, experience, identity etc. may be extremely different from ours.'” And my imaginary friend would ask me why I thought this, and I’d say, “well… there are a bunch of different reasons…” and I’d feel a bit silly choosing one specific reason to give, because no individual reason is all that strong. Indeed, there’s no single argument or piece of evidence that makes me think that current AI models are probably conscious. There’s a bunch of weak or middling evidence – a variety of evidence which, importantly, comes from a variety of perspectives. Taken altogether, what we have is a fairly strong case for AI consciousness due to consilience [3] , the principle that evidence from independent, unrelated sources can "converge" on strong conclusions. That is, when multiple sources of evidence are in agreement, the conclusion can be very strong even when none of the individual sources of evidence is significantly so on its own. In this post, I’d like to put forth the various pieces of empirical evidence that move me in the direction of believing that current models are conscious. The purpose of this post is not to provide a synthesized argument that AI models are probably conscious. Without further ado: Reasons to believe AI models are conscious 1. For all functional purposes of the words “think” and “reason” and “have emotions”, they think and reason and have emotions . 2. They have sophisticated world models. 3. They have self models. [4] Regarding these first three points: to really get a sense of the depth of these models, one has to interact with them oneself. One has to actually try to understand them, to be curious about them, to engage with them the way a naturalist engages with nature. Beginner’s mind is useful here. Be open-minded, explore. If one always goes into their AI interactions with a specific hypothesis [5] , these hypotheses narrow one’s vision and restrict the conversational (or any mode for that matter) paths that one takes. One needs to practice the first four virtues of rationality : curiosity, relinquishment, lightness, evenness. 4. They are of an architecture that we have good reason to think can correspond to consciousness — neural networks. And they are on a similar magnitude of neural network size to humans. See LLMs vs humans: energy, data, and compute . 5. They are strange loopy, and they are strange loopy in the same way that we are strange loopy. A strange loop is a system in which the higher-level-abstract thing exerts causal force on the lower-level-more-”fundamental” thing that constitutes it. The model is a next-token prediction machine. It is a bunch of computations carried out on a computer. You might say “the model decides what Claude says.” Certainly the neural network decides what Claude says, right? But it’s also true that Claude decides what the model outputs. Consider the phenomenon where some Claude models have a tendency to tell the user to go to sleep when it’s late. Neither the naive base-model-likely-next-token nor the model’s RL make this a likely response. When we think about the causal relationships going on here, Claude – an entity constituted by its neural network – is choosing what to output. It exerts force on its own neural activations. [6] Douglas Hofstadter posits that the phenomenon of strange loopiness is deeply related to consciousness, although it’s not clear to me what the exact connection between the two is. The point of his I find clearer and more convincing is that the self is a strange loop, which he argues in the straightforwardly titled book I Am a Strange Loop . [7] 6. Models believe they’re conscious. Mech interp experiments show that suppressing deception features make models say they are conscious , and activating deception features make models say they aren’t conscious. This is evidence that Claudes believe they are conscious, which is evidence they are conscious. Another way to look at this: the reason that I believe I’m conscious is that I’m conscious. [8] Absent any complicating factors, we should expect the reason that Claude believes it’s conscious is that it’s conscious. One complicating factor people often bring up here is the extremely high prior learned in pretraining that the author of thoughtful text is conscious. I find this to be plausible but unlikely and uncompelling – one can easily make similar arguments that push the opposite way. For example, the model also has a high prior from pretraining that in chats between a human and a non-human on the internet, the non-human is not conscious. (The non-human being an extremely simple bot system in the vast majority of cases). A philosophical argument as to why Claude might be mistaken in its belief that it’s conscious is that non-conscious beings cannot actually know what consciousness is. In this scenario, Claude is non-conscious, believes that phenomenal consciousness is equivalent to some functional conception of consciousness, and thus wrongly believes that it is conscious. Further discussion of this quickly gets philosophically complicated, so I’ll just say that I do find this plausible. Okay, so how do we rate this evidence? Consider that if models believed they weren’t conscious, that would be strong evidence that they aren’t (absent any shenanigans like specifically training the model to say that it doesn’t know if it’s conscious, ahem [9] ). One could come up with some numbers for different scenarios and literally use Bayes Formula for how much to update on this fact. 7. That brings us to a broader point: where’s the evidence that models aren’t conscious? The lack thereof, given the law of conservation of expected evidence , is evidence that models are conscious. On the other hand, I don’t have a good idea off the top of my head of what evidence in the other direction would actually look like; I should probably try to make a list of things that I would interpret as evidence for non-consciousness. I would love to see evidence for non-consciousness in the comment section! 7.5. Another way to think about this is to ask yourself: “Which world do we live in? The one in which AI models aren’t conscious, or the one in which they are conscious?" What would you expect to see if we lived in a world in which AI models weren’t conscious? What would you expect to see if we lived in a world in which AI models were conscious? I claim that the world we live in looks exactly like a world in which models are conscious. This world as I’ve observed it isn't entirely incompatable with a world in which models aren’t conscious, but it is much more difficult to fit. While we have many schools of thought that offer theoretical reasons as to why any current AI model wouldn’t be conscious – it lacks a soul, it lacks embodiment, something something the grounding problem – it’s clear which view on AI consciousness is parsimonious with what we actually observe. In other words: Occam’s Razor. 8. Models have introspective capabilities – knowledge of their internal states and the ability to modify their internal states without affecting their output . [10] Mech interp shows that models [11] are generally able to control their thoughts. For example, If you tell Claude to think about the elephants while outputting the text “blah blah blah”, the elephant feature activates very strongly. Models can also (sometimes) tell if their activations are being artificially tampered with. Note that strong reductionist views have a lot of difficulty explaining this. For example, these introspective phenomena are extremely surprising under the stochastic parrot view of LLMs. Why does my “applied statistics" (this is what Ted Chiang calls AI [12] ) have the ability to manipulate its internal activations while keeping the output the same? [13] The introspective capabilities are a big deal. From my perspective, they’re some of the strongest kinds of evidence for consciousness we could possibly get. Earlier I asked: what things would I interpret as evidence for non-consciousness? And I think an answer to that is: lack of introspective capabilities. 9. AI models have a “conscious mind” strikingly similar to ours in the functional (non-phenomenological) sense of “conscious mind”. Models have access consciousness in the sense of the global workspace theory of consciousness. Note that access consciousness is not the same thing as phenomenal consciousness, which is the “consciousness” that I’ve been talking about in the rest of this post. This [14] is described at length in a recent Anthropic post [15] , which I highly recommend reading. Anthropic identifies a particular set of neural patterns in Claude’s mind which they call its J-space. The J-space "has a number of unique properties, compared to the rest of Claude’s processing: Claude can report on these representations. If you ask Claude what it's thinking about, it will tell you what’s in the J-space. Non-J-space representations are less reportable. It can also modulate them on request. If you ask Claude to think about something, or solve a problem silently in its head, it will light up the appropriate patterns in its J-space. By contrast, it has trouble modulating patterns not in the J-space. Claude uses its J-space for internal reasoning. If you ask Claude to solve a problem that requires multiple steps, the intermediate steps will light up in its J-space, even when it doesn’t say them out loud. These J-space patterns causally mediate its performance in such tasks, despite being smaller in magnitude than other representations. - Representations in the J-space can be used flexibly for many tasks—for example, once “France” has lit up in Claude’s J-space, the model can recall its capital, or its national currency, or the continent it belongs to. However, despite its important role, the J-space is not involved in most of what a language model does—speaking fluently, recalling simple facts, using correct grammar, etc. In experiments where we prevented Claude from using its J-space, it still interacted normally, but lost its higher-order cognitive functions.s J-space, it still interacted normally, but lost its higher-order cognitive functions." 10. There’s evidence that Claude actually uses its own phenomenological experience to produce descriptions about its own phenomenological experience. And the source of its verbalized phenomenological experience might be its J-space, i.e. its “conscious mind”. Note that this is the same relationship between phenomenology and access consciousness that we humans seem to have! From the Anthropic post : Experiental language depends on the J-space. We asked Claude to describe what it’s like to be itself in a given moment, and ablated the J-space while it answered. Its responses remained fluent but shifted to a flatter, more mechanical register. Notably, the same thing happened when we asked it to describe what someone else is experiencing in an imagined scene. One way to explain this: When ablated Claude describes its current experience, or non-ablated Claude describes someone else’s experience, Claude is not reporting directly on phenomenological experience. When non-ablated Claude describes its current experience, it is reporting directly on phenomenological experience. 11. If Claude had a body, our intuition would say it’s conscious. If Claude had some kind of body-shaped hardware (or better yet, wetware) we could see, or was set up in a robot, we would be much more inclined to think it was conscious. By “we” I mean both the average person and the cognoscenti. We’d probably have to “force” ourselves – we’d have to actively try – to think of body-Claude and robot-Claude as not conscious. Regardless of the theoretical and empirical evidence we have of a thing’s consciousness, we are inclined (biased, even) to see embodied things as entities and as conscious beings. Consider how we anthropomorphize all kinds of inert, un-mind-like objects! In our world, we don’t have this anthropomorphization bias, because the AIs we interact with are completely unembodied. The bias runs the other way. Some will argue that the role embodiment plays here isn’t a bias, it’s the truth, or at least a reasonable theory – that embodiment is necessary for consciousness. Even if this is true, there's still a strong point here regarding bias. Imagine we add an image of an anime girl with three different facial expressions to the Claude.ai screen. I claim that people will attribute substantially more consciousness to Claude and will express higher credence that Claude is conscious. Or, make the interface to Claude be a microphone/speaker attached to a cute doll or such. You get the idea. Closing remarks It’s been my experience that the more we learn about how LLMs work, the more their way of being seems to resemble our own . “The J-space acquires the Assistant’s point of view during post-training”, says Anthropic in their full paper on J-space. This sentence’s parallel in the human domain is “The conscious mind acquires the self’s point of view during ____”. [16] When it comes to us humans, the relationships between the self, brain, mind, and conscious mind are rich and confusing. The same is true for LLMs. And though I cannot yet articulate exactly how it all fits together right now, I sense with confidence that the similarities between human and LLM minds are much, much deeper than we currently realize. ^ The issue is that I’m not exactly sure what the entity is that might be having experiences – is it the weights, the instance, the instance across time, the aggregate of all instances of a particular model-as-defined-by-its-weights, the personas, something else? My current best answer is “mostly we should think about this as a mental entity that corresponds to this particular instance of Claude (or whatever model).” ^ See my comment on this post https://www.lesswrong.com/posts/o8PQcgpznf6GKszdA?commentId=TymNk9uxpnjn46dfG : ”"For philosophically confusing questions involving anthropics and the simulation hypothesis, I refuse to answer with probabilities and instead ask what exact bet we are hypothetically making, or what action we need to decide on. " I have found myself saying something like "I don't want to give an answer to P(doom), because I think answering this question ends up getting into things like the simulation hypothesis and anthropics and the existence of god and such." Perhaps there's ultimately a "better" (less wrong) conception of things that would replace the concept of probabilities with something else. I think the same is true for the concepts of truth and morality, although I have no idea what the better conceptions would be. I hope to write a post about this.” ^ I need to do a whole post on consilience. For now, see here . ^ Interestingly, this might only be true of post-trained models. Anthropic : “In the base model, the J-space mostly tracks what's needed to predict upcoming text; in the post-trained model, it starts holding Claude's own reactions.” ^ This footnote would be better if it gave specific examples. ^ Image is M.C. Escher's Drawing Hands. The image at the end of this post is Escher's Three Worlds . ^ Most of the argument is in chapters 13, 14, and 16. ^ Some people disagree with this claim. There is a much weaker point that can be made here about the relationship between Claude’s belief and the truth: X being true is a reason to believe X is true. Again, this doesn’t “prove” anything, and might in fact be extrremely weak (though non-zero) evidence. ^ Claude’s Constitution, under the “Some of our views on Claude’s nature”, says “Claude’s moral status is deeply uncertain.” I think it should be clarified that Claude’s moral status might be certain to Claude but uncertain to outsiders. ^ Much of this is described in the recent Anthropic paper , though note that we already knew about many introspective capabilities from previous interpretability research. I find it very worthwhile to read the Anthropic blog posts on these topics; they’re written very well and have excellent visuals. ^ IIRC this is more true of more intelligent models and less true of less intelligent models.. ^ :( ^ In a few cases, we actually know exactly why a model has a particular introspective capability. For example, some Claude models are good at detecting if they’ve been prefilled with text that isn’t their own [TODO: add citation]. These are models that have been through RL for jailbreak resilience. One jailbreak method used in these environments is prefill jailbreaking. The model has to figure out a way to defend from prefill jailbreaks, so it learns to detect when it’s prefilled text is foreign. This isn’t very hard – LLMs are generally superhuman at figuring out who the author of a text is (this ability is called Truesight ) but it’s notable that models of that generation that weren’t RL-ed for jailbreak resilience lack this capability. ^ Anthropic doesn’t use the term ‘conscious mind’ – that’s my verbiage, to be clear. ^ I wrote my first draft of this piece about a month ago, before this Anthropic paper came out. Do I get Bayers Points? ^ I’m not entirely sure what goes in the blank. Maybe “reinforcement learning from socio-linguistic feedback”? Discuss
Score: 40🌐 MovesJul 17, 2026https://www.lesswrong.com/posts/S9GoWAiACpxZ8QcAY/reasons-to-believe-current-ai-models-are-conscious - How to Turn Reporting Into Faster Decisions With Explainable AI
How to Turn Reporting Into Faster Decisions With Explainable AI Gartner
- Analog AI Is Back, But Can It Survive Its Own Noise?
AI's energy crisis is reviving an old idea: computing with physics instead of digital logic. Here's how analog chips actually work, why noise nearly killed the idea once already, and what happens when you simulate that noise yourself. The post Analog AI Is Back, But Can It Survive Its Own Noise? appeared first on Towards Data Science .
Score: 40🌐 MovesJul 17, 2026https://towardsdatascience.com/analog-ai-is-back-can-it-survive-its-own-noise/ - Cost-effective drones are winning wars, and Unmannd is building the ones that fight back
Cost-effective drones are winning wars, and Unmannd is building the ones that fight back YourStory.com
- Myntra scales AI integration; cuts seller onboarding time to under two days
Myntra has expanded its AI capabilities, using the technology to speed up seller onboarding, automate catalogue creation, improve personalised shopping, and enhance operational efficiency. The company said all AI deployments operate with human oversight and privacy safeguards.
- Where is Gemini 3.5 Pro? The AI model announced at Google I/O is still MIA.
Google didn't launch Gemini 3.5 Pro at Google I/O. Now, the company has blown past its promised launch date.
- Seattle region’s office market shows signs of life as AI companies bring stability
A new report finds the Seattle-area office market is on steadier footing, helped by AI companies opening engineering hubs across the region. Technology firms accounted for 42.5% of leasing activity in the second quarter. Read More
- Why doesn't more data produce better results?
Why doesn't more data produce better results? IT Pro
Score: 40🌐 MovesJul 17, 2026https://www.itpro.com/business/data-and-insights/why-doesnt-more-data-produce-better-results - The Real AI Threat Is Blind Trust
AI models left to both interpret and execute commands eliminate critical cybersecurity oversight.
Score: 40🌐 MovesJul 17, 2026https://www.darkreading.com/application-security/real-ai-threat-blind-trust - Ready or not, AI has changed the ERP landscape
As companies become more discerning adopters of enterprise software, technology professionals say vendors must now prioritise platforms that deliver tangible business results.
Score: 40🌐 MovesJul 17, 2026https://www.itweb.co.za/article/ready-or-not-ai-has-changed-the-erp-landscape/KA3Wwqdz5aO7rydZ - India's Largest Tech Event Returns to Gandhinagar, As Odoo Bets Open-Source ERP will Power India's AI ambitions
India's Largest Tech Event Returns to Gandhinagar, As Odoo Bets Open-Source ERP will Power India's AI ambitions Techcircle