AI News Archive: August 10, 2026 — Part 5
Sourced from 500+ daily AI sources, scored by relevance.
- Emirates NBD joins Dubai fund to bring more AI and FinTech solutions to banking
Emirates NBD joins Dubai fund to bring more AI and FinTech solutions to banking Gulf News
- Zuckerberg warns against centralizing AI power
In a 6,500 word essay, Zuckerberg detailed his vision for AI.
- Opinion | The Next ‘Lab Leak’ Could Be AI
Recent sandbox breaches demonstrate the need for federal guardrails.
Score: 38🌐 MovesAug 10, 2026https://www.wsj.com/opinion/the-next-lab-leak-could-be-ai-a774e763?mod=rss_Technology - Discovered Materials is playing AI whack-a-mole to hunt cooler chips
Discovered Materials raised $9 million to fund the hunt for more novel materials to build more efficient chips.
Score: 38🌐 MovesAug 10, 2026https://techcrunch.com/2026/08/10/discovered-materials-is-playing-ai-whack-a-mole-to-hunt-cooler-chips/ - State chatbot laws could create regulatory patchwork, new report warns
Researchers claimed inconsistent state definitions and requirements could make it harder for developers to comply, while diverting attention from safeguards that directly address documented risks to children.
Score: 38🌐 MovesAug 10, 2026https://statescoop.com/state-chatbot-laws-could-create-regulatory-patchwork-new-report-warns/ - The Roboguard Revolution is Short-Circuiting
Knightscope and other robotics companies are rethinking automated security following canceled contracts. One pivot? Human guards.
- Globant launches new AI consultancy marketplace for AI services | ChannelPro
Globant launches new AI consultancy marketplace for AI services | ChannelPro IT Pro
Score: 38🌐 MovesAug 10, 2026https://www.itpro.com/business/business-strategy/globant-launches-new-ai-consultancy-marketplace-for-ai-services - Rackspace beats expectations but its losses pile up amid aggressive AI pivot
Shares of Rackspace Technology Inc. were trading lower after-hours today, despite an encouraging earnings and revenue beat in its second-quarter financial results. The San Antonio-based company reported earnings before certain costs such as stock compensation of eight cents per share, just ahead of Wall Street’s target of nine cents per share. Revenue for the period […] The post Rackspace beats expectations but its losses pile up amid aggressive AI pivot appeared first on SiliconANGLE .
Score: 38🌐 MovesAug 10, 2026https://siliconangle.com/2026/08/10/rackspace-beats-expectations-losses-pile-amid-aggressive-ai-pivot/ - UAE launches AI-powered higher education platform to boost university performance
UAE launches AI-powered higher education platform to boost university performance Gulf News
- Why detecting deepfakes is structurally harder than building them, and what it will take to close the gap
By Ankush Tiwari, Founder and CEO, pi-labs Somewhere in India this month, a man is watching a video of himself say something he has never said. The voice is his. […] The post Why detecting deepfakes is structurally harder than building them, and what it will take to close the gap appeared first on Express Computer .
- How to ground Genie Agents in both structured data and documents without losing governance
Building an agent to automate simple business tasks can be easy. But creating one...
- Small Language Models Are Eating Your LLM Bill: A Fine-Tuning Cost Breakdown
When a fine-tuned 7B model beats a frontier API — reportedly a 20x–100x cost difference on narrow, high-volume tasks. Continue reading on Towards AI »
- Former ByteDance robotics head reportedly joins Xiaomi
Former ByteDance robotics team head Kong Tao has reportedly joined Xiaomi, where he now leads a team developing foundation models for robots. Several people familiar with the matter said Kong joined Xiaomi in 2025 and brought a number of former ByteDance colleagues with him. Xiaomi’s robotics division is said to have about 200 employees, while […]
Score: 38🌐 MovesAug 10, 2026https://technode.com/2026/08/10/former-bytedance-robotics-head-reportedly-joins-xiaomi/ - Innocent shopper kicked out of Sainsbury’s after facial recognition error
Innocent shopper kicked out of Sainsbury’s after facial recognition error The Telegraph
Score: 38🌐 MovesAug 10, 2026https://www.telegraph.co.uk/news/2026/08/10/innocent-shopper-sainsburys-after-facial-recognition/ - Senior UK detective under investigation for alleged misuse of AI
Police watchdog says officer was referred to them by Derbyshire Constabulary following ‘internal review’
Score: 38🌐 MovesAug 10, 2026https://www.ft.com/content/28b389f5-5f74-4993-a9b3-edd0c500d49a?syn-25a6b1a6=1 - Impact-resistant, autonomous robots inspired by tensegrity architecture
Nature Machine Intelligence, Published online: 10 August 2026; doi:10.1038/s42256-026-01280-2 Johnson et al. demonstrate an autonomous three-bar tensegrity robot capable of robust locomotion across varied terrains even after extreme impacts, including a 5.7-m drop onto asphalt.
- Zuckerberg Calls U.S. to Reduce Friction Around Open-Source
Zuckerberg Calls U.S. to Reduce Friction Around Open-Source The Information
Score: 38🌐 MovesAug 10, 2026https://www.theinformation.com/briefings/zuckerberg-calls-u-s-reduce-friction-around-open-source - Inflation Data Will Be the Real Test for This AI Stock Rally
Inflation Data Will Be the Real Test for This AI Stock Rally Barron's
Score: 38🌐 MovesAug 10, 2026https://www.barrons.com/articles/stock-market-things-to-know-today-40709717 - Leaked memo: The Arena Group rebrands to Paradium.AI in pivot to 'publicly traded AI company'
Leaked memo: The Arena Group rebrands to Paradium.AI in pivot to 'publicly traded AI company' Business Insider
Score: 38🌐 MovesAug 10, 2026https://www.businessinsider.com/the-arena-group-rebrands-paradium-ai-ceo-memo-2026-8 - AI agents, IRL socializing, and vibe coding: Where top consumer investors are placing bets
AI agents, IRL socializing, and vibe coding: Where top consumer investors are placing bets Business Insider
Score: 38🌐 MovesAug 10, 2026https://www.businessinsider.com/next-big-bets-consumer-tech-according-to-top-vcs-2026-8 - AI professors are negotiating the new realities of academic research
This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. Last week, I headed 30 miles south of San Francisco to a hotel in Mountain View, California, to join some of the most accomplished, and some of the most promising, AI…
- The AI reckoning every CIO saw coming (and still wasn’t ready for)
Earlier this year, the National Bureau of Economic Research released survey results from over 6,000 U.S. leaders showing that while AI adoption is widespread at 69%, we’re seeing little to no impact on productivity. Anecdotally, we’ve seen leaders from top companies echo that refrain. It’s the reckoning many CIOs, CTOs and COOs are navigating as we enter the last half of the year. The most humbling part is knowing it’s a management problem we created by treating AI like it was exempt from the rules we apply to every other enterprise tool. Part of this has to do with how AI entered the market. The tools that sparked its mainstream adoption arrived as consumer products before enterprises had governance frameworks to absorb them. Enterprises were left playing catch-up as they grappled with IP and data security concerns, inadvertently fueling shadow AI as employees leveraged these tools to get ahead and eventually, keep pace, at work. What this created was a sense of entitlement that is challenging to unravel. Like the internet writ large, employees have grown to expect unlimited access, and organizations played along. But this idea warrants a pause. When did we last roll out Salesforce to everyone who asked without a use case? AI got a pass because it felt different. In truth, it isn’t. It’s another tool that enterprises need to manage. Three levels every information and technology leader has to solve It’s helpful to look at this as a three-level evolution framework. Level one is adoption — are people actually using it well? Level two is budget control — what are we spending and on what? Level three is justification — can we demonstrate the return? Most companies are still at level one. Deloitte reported in their 2026 State of AI in the Enterprise report that only 25% of respondents have moved 40% or more of their AI experiments into production to date. The minority that have moved pilots to production are grappling with the budget and trying to figure out how to justify the costs and quantify the gains. The problem is, you can’t prove what you didn’t have to hire because of AI. There is no parallel universe where you can walk into the CEO’s office and say I need five more people in finance, but in this universe, with AI, I didn’t. The organizations that wait for a clean ROI model before making any decisions will spend themselves into the trough of disillusionment before they find one. The smarter move is to start treating it as a discipline you build. What managing AI like a tool actually looks like To some, governance sounds like restriction. But the discipline is more about matching the right tool to the right use case, and making the sanctioned path easier than the workaround. Take shadow AI . The instinct is to lock things down. But when employees start building internal apps with company data and hosting them on free public platforms, the answer isn’t another policy. By the time the policy is written, the data is already public. Instead, you need to build an internal alternative that does the same thing without the exposure. Give people a path. If you don’t, they build their own, and you won’t know about it until something goes wrong. The same principle holds for conflicting data. Two departments pulling AI-generated recommendations from the same underlying data and arriving at different conclusions isn’t an AI problem. It’s a data and definitions problem. AI just made it impossible to ignore. Say marketing claims they brought $50 million in the pipeline, and sales claim they brought $50 million as well.But the company actually has $75 million in pipeline. Someone is counting the same deals twice under different definitions. The CIO’s job is to enforce one source of truth. If your dashboard doesn’t match the authoritative one, your dashboard is wrong. That’s the only way the organization can function. And it applies to cost, too. Not every workflow needs the most expensive model. Not every employee needs full AI access. If someone is using a top-tier model to summarize email because nobody told them there was a cheaper option that does the job, that’s a gap that CIOs need to address. The CIO’s job is to build the layer that makes the right choice the obvious one, and to provide sanctioned alternatives so employees aren’t left building their own. That’s what actually reduces shadow AI, conflicting data and runaway spend: Alternatives, visibility and a single source of truth. The CIOs getting real value from AI right now aren’t the ones who said yes to everything. They’re the ones who asked the same questions they’d ask about any other enterprise investment: What does it do, who actually needs it and what are we getting back? AI is a remarkable tool. It’s also just a tool. It doesn’t exempt you from the management discipline you apply to every other system in your stack. We didn’t roll out Salesforce to everyone who asked without a use case. We shouldn’t have done it with AI either, and the organizations that did are now living with the consequences: Six-figure token bills, shadow apps on public URLs, dashboards that contradict each other and a CEO asking what exactly he got for the investment. The answer to that question is available. But only if you built the infrastructure to find it.
Score: 37🌐 MovesAug 10, 2026https://www.cio.com/article/4206733/the-ai-reckoning-every-cio-saw-coming-and-still-wasnt-ready-for.html - Digital Science connects AI agents to world-leading research data with new Dimensions MCP servers
Digital Science connects AI agents to world-leading research data with new Dimensions MCP servers EurekAlert!
- J.P. Morgan lifts 2026-end target for S&P 500 to 8,000 on AI, earnings strength
J.P. Morgan lifts 2026-end target for S&P 500 to 8,000 on AI, earnings strength Reuters
Score: 36🌐 MovesAug 10, 2026https://www.reuters.com/business/jp-morgan-lifts-2026-end-target-sp-500-8000-ai-earnings-strength-2026-08-10/ - KIST develops neuromorphic AI training technique to usher in the era of low-power AI
KIST develops neuromorphic AI training technique to usher in the era of low-power AI EurekAlert!
- Ford’s new AI assistant can check your fuel levels and tire pressure
The assistant is rolling out to the Ford mobile app now. And by 2027, it will be available through the vehicle itself.
Score: 36🌐 MovesAug 10, 2026https://www.theverge.com/transportation/976748/ford-ai-assistant-mobile-app - If software and agentic AI are key to mission success, accelerate them to the front lines
If software and agentic AI are key to mission success, accelerate them to the front lines Breaking Defense
- UAE Islamic Bank ruya Partners With Magure to Build an AI-Native Banking Model
UAE Islamic Bank ruya Partners With Magure to Build an AI-Native Banking Model Entrepreneur Middle East
- Prompt Caching vs. Fine-Tuning: A Cost and Latency Decision Framework
In this article, you will learn how prompt caching and fine-tuning differ as strategies for reducing cost and latency in agentic AI systems, and how...
Score: 36🌐 MovesAug 10, 2026https://machinelearningmastery.com/prompt-caching-vs-fine-tuning-a-cost-and-latency-decision-framework/ - Adesso and Hitachi Digital Services partner to accelerate AI-led enterprise transformation
German technology consulting and IT services company Adesso SE and Hitachi Digital Services have entered into a strategic partnership to help enterprises accelerate transformation through AI, cloud modernisation, digital engineering and IT/OT convergence. The post Adesso and Hitachi Digital Services partner to accelerate AI-led enterprise transformation appeared first on Express Computer .
- Muse Code vs Claude Code vs Codex CLI: 7 architecture differences worth evaluating
Muse Code, Claude Code, and Codex CLI now match on features but diverge on orchestration, state, and security. Here’s the architecture… Continue reading on Towards AI »
- Ford’s new AI assistant saves you from googling that dashboard warning light
Ford's new AI assistant, live today in the Ford and Lincoln apps, answers real questions about your specific vehicle using live telemetry, with a full in-car version planned for 2027.
Score: 35🌐 MovesAug 10, 2026https://www.digitaltrends.com/cars/fords-new-ai-assistant-saves-you-from-googling-that-dashboard-warning-light/ - It Isn’t Too Late to Buy AI Hardware Boom, Analyst Says. Just Look at HPE Stock.
It Isn’t Too Late to Buy AI Hardware Boom, Analyst Says. Just Look at HPE Stock. Barron's
Score: 35🌐 MovesAug 10, 2026https://www.barrons.com/articles/hewlett-packard-enterprise-hpe-stock-ai-price-43f20b6f - How In-House Compliance Teams use AI to Stay Ahead of Regulatory Change
Explores how AI helps compliance teams anticipate and adapt to regulatory shifts.
- Old OCR text cripples language model training, and FineBooks wants to fix that at scale
The FineBooks project from Hugging Face and EleutherAI tested 14 open-source OCR models on more than 2,000 historical book pages. The top model, dots.mocr, hits 97.6 percent character accuracy at under two dollars per thousand pages. That's good enough for AI training data, but not yet for scholarly transcriptions, the team says. The article Old OCR text cripples language model training, and FineBooks wants to fix that at scale appeared first on The Decoder .
Score: 35🌐 MovesAug 10, 2026https://the-decoder.com/old-ocr-text-cripples-language-model-training-and-finebooks-wants-to-fix-that-at-scale/ - US does its robotics industry no favours by fencing it off from China
Concession speeches follow a formula: the vocabulary of defiance, thanks to the faithful, a promise that the fight goes on. Last month, Brendan Carr, chairman of the US Federal Communications Commission (FCC), essentially delivered one on behalf of American robotics. Acting on findings from a White House task force, Carr added new foreign-made humanoid robots, quadrupeds and power inverters to the agency’s Covered List, denying them the authorisation nearly every electronic device requires to be...
- What happens when medical students rely on AI – and never develop their own judgment? | Simar Bajaj and Joseph Sakran
AI’s danger isn’t just in experts losing the ability to reason. It’s that trainees may never learn how to do so in the first place In healthcare, there’s growing concern over doctors becoming less clinically adept as they increasingly rely on AI tools. But what about the trainees – medical students, residents and fellows – who are using these tools before they have built their own clinical judgment? The idea of deskilling implies that someone possessed an ability and then lost it. Here, the danger is not just deskilling but never-skilling. Although a doctor who has forgotten how to reason is recoverable, one who never learned how may not be. OpenEvidence, essentially an AI chatbot for clinicians, has given this concern its most concrete form. About two-thirds of US doctors actively use OpenEvidence, asking about puzzling symptoms, drug interactions and clinical guidelines, getting responses within seconds, anchored in the latest research. Trainees, unsurprisingly, have also begun to use this AI tool in many of the same ways – but at a far more formative stage. Continue reading...
Score: 35🌐 MovesAug 10, 2026https://www.theguardian.com/commentisfree/2026/aug/10/ai-medical-students-judgment - AI security emerging as separate budget line for Indian enterprises: Palo Alto Networks’ Swapna Bapat
Palo Alto Networks’ Swapna Bapat said awareness of the risks from unsecured AI is growing among chief information security officers. “If I have to use AI, I have a budget for using AI, then I have a budget to secure AI as well,” she said.
- Todd Boehly’s group rolls out AI to Chelsea and A24
Eldridge will embed technology in companies from Chelsea Football Club to film studio A24 after acquiring 50% of Sudolabs
Score: 35🌐 MovesAug 10, 2026https://www.ft.com/content/1b3493a5-2b0b-4a3e-86e0-76b7f2de373a?syn-25a6b1a6=1 - Robot Recycler Salvages Parts from Broken Machines
This system is getting the automated circular economy rolling
- Google’s search monopoly is becoming an AI monopoly
Google's search crawler now feeds its artificial intelligence models, creating an unfair advantage. This strategy accelerates automated internet traffic, surpassing human activity online. Regulators and industry players are pushing for changes to this practice. Google will now allow websites to opt out of AI data usage without search ranking penalties. This shift aims to ensure original content creators are still rewarded for their work.
- AI is changing more than young bankers' workflows. It's creating a whole new career path.
AI is changing more than young bankers' workflows. It's creating a whole new career path. Business Insider
Score: 35🌐 MovesAug 10, 2026https://www.businessinsider.com/ai-startup-rogo-wall-street-careers-junior-bankers-2026-8 - Open vs. closed source AI: Is the future of AI free or behind a paywall? | Fortune AI Playbook
Open vs. closed source AI: Is the future of AI free or behind a paywall? | Fortune AI Playbook Fortune
- The Morning Download: Meta Shares Glimmer of Always-On AI Future
Meta’s new model puts always-on AI agents front and center
- AI-powered system could help save lives by predicting how modern house fires behave
AI-powered system could help save lives by predicting how modern house fires behave EurekAlert!
- The limits of physics AI: where Siemens says the human stays in charge
Physics AI can now explore thousands of design variations in the time it would take a traditional simulation to chew through a handful of them. Precisely up to 1,000 times faster, according to Siemens. What it cannot do is sign off a safety-critical part. On that, the technology has a firm limit, and Sam Mahalingam, […] The post The limits of physics AI: where Siemens says the human stays in charge appeared first on AI News .
Score: 35🌐 MovesAug 10, 2026https://www.artificialintelligence-news.com/news/siemens-physics-ai-simulation-human-oversight/ - Can AI Command Earth-to-Orbit Operations?
The aerospace and defense sector is facing a confluence of geopolitical instability, rapid technological advances, evolving security requirements, and complex global supply chains. The post Can AI Command Earth-to-Orbit Operations? appeared first on EE Times .
- 49ers coach Kyle Shanahan says his Tesla was on Autopilot before he crashed, but it's 'always your fault'
49ers coach Kyle Shanahan says his Tesla was on Autopilot before he crashed, but it's 'always your fault' Business Insider
Score: 34🌐 MovesAug 10, 2026https://www.businessinsider.com/kyle-shanahan-49ers-coach-tesla-autopilot-crash-2026-8 - I Built a Production RAG System for Indian Law and Refused to Trust It Until I Measured It
How I turned a stack of legal PDFs into a chatbot that cites the exact section of the law, and why the evaluation scores, not the demo, were the real project. Ask a lawyer in India a simple question “what’s the punishment for cheating?” and you’ll get an answer in thirty seconds. Ask that same question as an ordinary citizen, and you’re staring at a 300-page PDF written in a register designed to keep you out. Most people can’t afford the thirty seconds of a lawyer’s time. So they guess, or they trust a stranger, or they do nothing. That gap is the reason I built Lexora: a chat app where you ask a question about Indian criminal law in plain language and get an answer grounded in the actual law , with the exact section cited BNS · Sec 318 · active so you can verify it yourself. This is the story of how it was built, but more honestly, it’s the story of how I learned not to trust a RAG system that looks like it works and what it took to measure whether it actually did. If you take one thing from this article, let it be this: the demo is a liar, and the evaluation harness is the only thing that tells you the truth. What Lexora actually is Before the internals, the product in three sentences: You ask a legal question. Follow-ups work it remembers the conversation. Every answer carries structured citation chips (e.g. BNS · Sec 103 · active), and each one is validated against the retrieved source , so the model physically cannot invent a citation. It’s a real account-based product: signup/login, saved conversations, history. The corpus is India’s criminal law both the new codes that took effect in 2024 (BNS, BNSS, BSA) and the repealed ones they replaced (IPC, CrPC, IEA), kept for old-vs-new context. That “old vs new” wrinkle matters more than it sounds, and it shaped a lot of the retrieval design. Here’s the whole system at a glance: Now let’s open the box. The RAG brain: retrieval is 80% of the battle A legal question has a nasty property: the keywords matter and the meaning matters, and neither alone is enough. Someone typing “Section 420” needs an exact keyword hit. Someone typing “what if I lied to get money” needs semantic understanding to land on the same offence. So Lexora uses hybrid retrieval: ChromaDB for semantic search over BGE-large embeddings catches meaning. BM25 for keyword search catches exact section numbers and legal terms of art. Both are combined with LangChain’s EnsembleRetriever, then a cross-encoder reranker (bge-reranker-base) re-scores the merged candidates and keeps the best few. That reranking step is underrated. Embedding similarity gets you roughly relevant chunks; a cross-encoder actually reads the query and the chunk together and tells you which ones truly answer it. It’s the difference between “these are in the neighbourhood” and “this is the one.” The router: don’t search what you don’t need Remember the old-vs-new problem? If someone asks about a current offence, dredging up the repealed IPC section pollutes the context. So before retrieval, an LLM router reads a “relationship map” of the corpus and picks metadata filters which act, which status (active vs repealed) so the search runs over the right slice. It’s a small, cheap LLM call that makes every downstream step better. This is a pattern I’d reuse anywhere: let a model narrow the search space before you search. Generation you can’t lie with The answering model (GPT) doesn’t just return prose. It returns structured output: an answer, a list of citations, and an answer_found boolean. Then comes the part I'm proudest of the citation-validation loop: Every section the model cites is checked against the sections that were actually retrieved. If it cites something that isn’t in the context, the answer is rejected and regenerated. A legal assistant that hallucinates a section number is worse than useless it’s dangerous so this loop is non-negotiable. It also produced one of my favourite small bugs. Validation kept failing on correct answers. The cause: the model returned "Section 103" while the metadata stored "103". String equality said "different." The fix was a lesson I keep re-learning: def normalize_section(s): # "Section 103", "Sec. 103", "103" → "103" match = re.search(r"\d+", s or "") return match.group() if match else None # compare normalize_section(cited) against normalize_section(stored) Normalize before you compare. Half of “AI bugs” are really string-formatting bugs wearing a trench coat. Memory ≠ dumping history at the model Conversational memory has a trap. When a user asks “what about culpable homicide?” as a follow-up, you cannot send that raw string to the retriever BM25 and embeddings aren’t an LLM, they have no idea what “what about” refers to. So Lexora runs a query-contextualization step first: an LLM rewrites the follow-up into a standalone question (“what is the punishment for culpable homicide under the BNS?”) before retrieval. The full history still goes to the answering model, but the retriever gets a clean, self-contained query. The lesson: a plain chatbot can dump history at the LLM; a RAG chatbot has to clean the query for the retriever separately. Two different consumers, two different needs. The part nobody blogs about: measuring whether it works Here’s where most RAG tutorials end “look, it answered my question!” and where the actual engineering begins. I did not want to feel like Lexora worked. I wanted a number. So I hand-built a golden dataset of 160 question–answer pairs and ran the pipeline through RAGAS, which scores four things: Faithfulness is the answer actually supported by the retrieved context, or is the model making things up? Answer relevancy does the answer address the question that was asked? Context precision of what we retrieved, how much was actually relevant? Context recall of what we needed , how much did we retrieve? The first run was humbling. Faithfulness was okay, but relevancy came back nan, and precision and recall were mediocre. A demo I'd have happily shown off was, by the numbers, mediocre. Diagnosing from the scores, not from vibes Two things were wrong, and RAGAS pointed at both. 1. The nan was infrastructure, not quality. RAGAS's native embeddings class didn't implement embed_query, so relevancy silently failed to compute. Wrapping the model properly (LangchainEmbeddingsWrapper) fixed it. Worth stating plainly: a nan is not a bad score, it's a broken measurement and confusing the two will send you optimizing the wrong thing. 2. The mediocre retrieval was a chunking problem. My first chunker (RecursiveCharacterTextSplitter) was packing five or six unrelated legal sections into a single chunk. That does two terrible things at once: it dilutes the embedding (one vector trying to represent six offences) and it corrupts the metadata (which section is this chunk even about?). No amount of reranking saves you from bad chunks. Fixing it was not one clean move it was a lot of trial and error. I rewrote chunking to be one section per chunk (a lookahead regex splitting on section boundaries, with the definitions section special-cased), restructured how each chunk’s metadata was built, upgraded embeddings from MiniLM to BGE-large, and tightened the citation handling then re-ran the harness, read the scores, and did it again. And again. Each pass moved a different metric. By the end, the scores had moved where it mattered most faithfulness, the one that decides whether a legal answer can be trusted, reached 93%: Chunking quality dominates RAG quality and getting there is iterative, not a single insight. If your retrieval is bad, fix the chunks before you touch anything else, then measure, then fix again. The integrity lesson that almost fooled me Then a subtle, scary one. Some RAGAS runs came back suspiciously good until I read the logs and found OpenAI rate-limit timeouts silently dropping questions from the average. The hard questions were timing out, getting excluded, and inflating the score. My RAG wasn’t getting better; my evaluation was quietly grading only the easy questions. The fix (a RunConfig with sane max_workers and timeout) was trivial. The lesson was not: "the score went up" means nothing until you check what got excluded from the average. An evaluation you don't audit is just a more sophisticated way of lying to yourself. Wrapping the brain in a product A RAG pipeline in a notebook is a science project. Making it a product was its own arc. API (FastAPI). I wrapped the pipeline in a clean /ask endpoint Pydantic request/response models, HTTPException handling so no stack trace ever leaks to a user, structured citations in the response shaped for the frontend , not the raw RAGAS internals. Auth & data (Supabase). Postgres tables for profiles, conversations, messages, with Row-Level Security on. This produced the best bug of the whole project. Conversation inserts kept failing with an RLS violation even though I was using the service-role key that's supposed to bypass RLS. The culprit: the supabase-py client is a shared singleton, and calling auth methods on it leaked the user's JWT into subsequent table calls, so my "admin" writes were silently running as the limited user. The fix was two separate clients: supabase = create_client(URL, SERVICE_ROLE_KEY) # DB writes supabase_auth = create_client(URL, ANON_KEY) # auth only Know your library’s hidden state. A shared client with mutable auth is a landmine, and the error message (“RLS violation”) pointed nowhere near the real cause. Frontend (Next.js). Landing → auth → chat, in a dark/gold brand, fully mobile-responsive (the chat sidebar collapses into a slide-in drawer). One process habit paid off repeatedly: I built the UI against mock data shaped like the real API first , so wiring the backend later was a swap, not a rewrite. The deployment gauntlet Shipping is where the estimates go to die. Host selection was a live-fire exercise. HuggingFace Spaces made Docker a paid feature the week I tried it. Hetzner’s cheap ARM instances were sold out everywhere, and their x86 8GB was €35/mo my mental price list was months stale. I checked live prices and pivoted to a Contabo VPS (~€5/mo, 8GB), running Docker + a Caddy reverse proxy, with the ML models downloading from the HuggingFace hub at runtime into a cached volume and all secrets living in a .env on the server, never in git. Then the classic. On the deployed frontend, chat just… did nothing. The console showed a mixed-content block: the page (HTTPS) was calling /conversations (no trailing slash), FastAPI was issuing a 307 redirect to /conversations/ but building that redirect URL as http://. Why? Uvicorn was behind Caddy, which terminates TLS, so uvicorn genuinely thought it was serving plain HTTP and had no idea the original request was HTTPS. The fix is one flag: uvicorn app:app --proxy-headers --forwarded-allow-ips=* Now uvicorn trusts Caddy’s X-Forwarded-Proto: https header and builds correct URLs. Behind a reverse proxy, you must tell the app the original scheme, or every URL it generates is subtly, invisibly wrong. Hardening for the open internet. Once it was public, the bots arrived instantly a steady drizzle of requests probing /wp-config.php and /.env (all 404, all normal). I locked CORS to the frontend origin, added rate limiting (/ask at 10/min, auth at 5/min which is as much about protecting the OpenAI bill as protecting the server), and put indexes on the hot database columns. What I’d tell my past self Technical RAG quality lives or dies on chunking and retrieval. Measure it (RAGAS), don’t eyeball it. The retriever is not an LLM. Give it clean, standalone queries. Validate every citation against the retrieved context so the model can’t hallucinate a source. Normalize before you compare. "Section 103" and "103" are the same fact in two costumes. Behind a proxy, forward the scheme or debug mixed-content errors at midnight. Know your library’s hidden state (looking at you, shared Supabase client). Process Ship first, harden second and decide what not to build so v1 actually ships. Build UI against mock data shaped like the real API. Verify against reality a real browser, real logs and always check what your evaluation excluded , not just the headline number. Estimates go stale. Check live prices and availability before you commit. Where it stands Lexora is live at lexora.cvijay.dev a stranger can visit a URL, ask a legal question in plain English, and get back an answer grounded in the actual section of the law, with a citation they can check. That was the whole milestone, and it’s real. Next on the roadmap: response streaming (the biggest remaining UX win), migrating embeddings to cut hosting cost, and expanding beyond criminal law into more domains. But the part I’ll carry into the next project isn’t the stack. It’s the discipline: I stopped trusting the demo the day the evaluation harness told me it was lying. If you’re building RAG and you don’t have a number yet that’s the first thing to build, not the last. If you’re working on legal AI, retrieval evaluation, or just want to compare notes on shipping RAG to production, I’d love to hear from you. I Built a Production RAG System for Indian Law and Refused to Trust It Until I Measured It was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.
- A brief guide to AI-powered software development environments
A brief guide to AI-powered software development environments InfoWorld
Score: 34🌐 MovesAug 10, 2026https://www.infoworld.com/article/4206868/a-brief-guide-to-ai-powered-software-development-environments.html