AI News Archive: August 19, 2026 — Part 11
Sourced from 500+ daily AI sources, scored by relevance.
- UGC-Style AI Video Prompts: 12 Examples for DTC and Creator Ads
12 UGC-style AI video prompts for direct-to-consumer and creator ads.
- How to Prompt Realistic AI Videos: 12 Prompts That Don't Look AI-Generated
12 prompts to create realistic AI videos that appear natural.
- Future-proof your career in the age of AI: Leadership expert shares ‘human’ strategies to stand out
Future-proof your career in the age of AI: Leadership expert shares ‘human’ strategies to stand out The Straits Times
Score: 24🌐 MovesAug 19, 2026https://www.straitstimes.com/singapore/jobs/future-proof-career-crystal-lim-lange-sph-media-career-expo - Entity accuracy in speech-to-text: why word accuracy isn't enough
Explores why entity-level accuracy matters beyond overall word error rate in STT.
- NCI to launch two bachelor’s courses in AI and cybersecurity
The courses are expected to launch in September 2027, subject to approval. Read more: NCI to launch two bachelor’s courses in AI and cybersecurity
Score: 24🌐 MovesAug 19, 2026https://www.siliconrepublic.com/innovation/nci-to-launch-two-bachelors-courses-in-ai-and-cybersecurity - Humanoid robots showcase kickboxing skills at Beijing tech show
The World Robot Conference in Beijing provides an opportunity for China’s robot makers to display their innovations and convince investors of the real-world potential for the technology.
Score: 23🌐 MovesAug 19, 2026https://www.nbcnews.com/video/kickboxing-robots-showcase-skills-at-beijing-tech-show-268551237779 - Deepfakes, Executive Impersonation and Corporate Trust: Andrea Baggio Details ReputationUP’s Response
Deepfakes, Executive Impersonation and Corporate Trust: Andrea Baggio Details ReputationUP’s Response USA Today
- AI vs. Automation: The Shift That’s Changing Agency Workflows
If you’ve evaluated new technology recently, you’ve likely noticed how quickly “AI-powered” capabilities become part of the conversation. From automated emails to workflow triggers to renewal reminders, many capabilities are grouped together. In practice, they serve different roles and are …
- This robot vacuum solves my kitchen stool problem
Hands-on with Mova’s new V70 Ultra Complete and its mopping arm that cleans hard-to-reach spots.
Score: 22🌐 MovesAug 19, 2026https://www.theverge.com/tech/981874/mova-v70-ultra-robot-vacuum-mop-arm-hands-on-review - Comparison of humans, GPT-4, and word embeddings for six basic emotions (IMAGE)
Comparison of humans, GPT-4, and word embeddings for six basic emotions (IMAGE) EurekAlert!
- The 6 AI-free Linux distros I recommend most - and why they're likely to stay that way
If AI isn't your jam, and you're looking for an operating system that doesn't (and won't) force it on you, look no further than these Linux distributions.
- More Wins, Less Admin: Why Startups Need Deal Management Software
Spend less time guessing and more time closing with these deal management tools.
Score: 22🌐 MovesAug 19, 2026https://www.salesforce.com/blog/small-business/deal-management-software-for-startups/ - What is AI Observability? A Complete Guide to Debugging and Monitoring Modern AI Systems at Scale
Your new AI product is live. Infra dashboards are all green. Latency is low, error rates are flat, and CPU/GPU utilization looks healthy. Yet your Slack channels are full of screenshots from users asking, “Why did it do this?” The same user goal and high-level prompt can still produce wildly different behavior. Sometimes the agent […] The post What is AI Observability? A Complete Guide to Debugging and Monitoring Modern AI Systems at Scale appeared first on Comet .
Score: 22🌐 MovesAug 19, 2026https://live-comet-marketing-site.pantheonsite.io/blog/what-is-ai-observability/ - Agrograde’s automated onion grading system installed at Pimpalgaon market
The machine processes up to 10 tonnes of onions per hour and cuts manual labour needs by 80%
- Exclusive: Krafton-backed Bobble AI enters insolvency process after failing to repay debt
Bobble.AI, the homegrown AI-powered communication and data intelligence startup, has entered insolvency proceedings. The National Company Law Tribunal (NCLT), New Delhi, has admitted an insolvency plea against Talent Unlimited Online Services Private Limited, the legal entity behind Bobble AI, and started the Corporate Insolvency Resolution Process (CIRP) against it, as per the filing sourced from the RoC. The order was passed on June 12 on a petition filed by Axis Trustee Services Limited, which was acting as debenture trustee for a group of investors. As per the filings, the tribunal has appointed Manish Agarwal as the Interim Resolution Professional (IRP) and declared a moratorium on the company. This means no legal action, asset sale, or recovery can happen against Bobble AI for now. According to the order and regulatory filings, Bobble AI had raised Rs 25 crore in March 2023 by issuing secured debentures to investors. The company started missing repayments from August 2025. With interest included, the total default stood at around Rs 5.77 crore. The security created against these debentures was also registered with the RoC, as per the filings. Bobble AI had argued that since talks with lenders were still going on, the default should not count yet. The tribunal did not accept this and said the company's own emails clearly admitted the debt and default. In an email response to Entrackr, Manish Agarwal, Resolution Professional of Bobble AI, clarified that the commencement of the Corporate Insolvency Resolution Process does not amount to the liquidation or closure of Bobble AI. According to him, the company continues to operate as a going concern during the resolution process, and its business operations, customer services, and ongoing commercial engagements continue in the ordinary course. Bobble.ai offers an AI-powered Indic keyboard supporting over 120 languages, with facial recognition for personalized GIFs and stickers. Integrated with WhatsApp and Facebook (Meta), the company monetizes through ads, data insights, subscriptions, and branded merchandise. The Gurugram-based company has raised around $35 million, including $26 million in a mix of primary and secondary capital from Krafton Inc in September 2022. Krafton controls a nearly 25% stake in the 12-year-old company. Last year, the company also laid off around 50 employees amid restructuring. Entrackr exclusively reported the development. Update at 5:22 PM, August 19 : The story has been updated with an official response from Bobble AI.
- The Night Before WRC 2026: Robots and Their People Are Still at Work at 10 PM
On August 18, the night before the 11th World Robot Conference opened in Beijing Yizhuang, workers were still installing booths, rehearsing boxing matches, and debugging robots until past 10 PM. The author chronicles how the event's big booths, real-scene demos, and focus on hands, perception, and data signal an industry shift from performance to value validation.
Score: 22🌐 MovesAug 19, 2026https://pandaily.com/world-robot-conference-2026-eve-beijing-yizhuang-robots-scenes-value-aug2026 - What is AI insurance?
What is AI insurance? IT Pro
Score: 22🌐 MovesAug 19, 2026https://www.itpro.com/technology/artificial-intelligence/what-is-ai-insurance - AI coding tools exposing a discipline gap in SA delivery teams
South African delivery teams have adopted AI coding tools faster than they've built the discipline to use them safely, says JustSolve.
Score: 22🌐 MovesAug 19, 2026https://www.itweb.co.za/article/ai-coding-tools-exposing-a-discipline-gap-in-sa-delivery-teams/LPp6VMrBNJZMDKQz - AI receptionist at GPs ‘can’t understand’ Yorkshire accents
The technology is said to struggle with ‘broad’ accents
Score: 21🌐 MovesAug 19, 2026https://www.the-independent.com/news/uk/home-news/ai-doctor-b3035818.html - Everyone’s Using This A.I. Dictation App That I Want to Murder With a Hammer
Wispr Flow is supposed to make writing effortless and magical. So I “wrote” this column with it.
- AI Centre boss Lee Hickin says ChatGPT is ‘like an abacus’ compared to AI’s practical uses
The NAIC boss says companies should stop hiding workplace use and build practical governance on your terms.
- Claude, ChatGPT, Gemini, and more are all included in this lifetime subscription, now just $99.99
1min.AI gives you lifetime access to dozens of AI models, including GPT, Gemini, and more
Score: 18🌐 MovesAug 19, 2026https://mashable.com/tech/aug-19-1minai-advanced-business-plan-lifetime-subscription - Building an evidence layer for AI
Organisations can prove that an employee used an AI system, but can they reconstruct what it was asked, what it produced and who relied on it?
Score: 18🌐 MovesAug 19, 2026https://www.itweb.co.za/article/building-an-evidence-layer-for-ai/kLgB17ezYAKM59N4 - A circuit prior in NN-bayes
Here are the slides of a talk Kaarel gave, presenting work with Dmitry establishing that (even arbitrarily overparametrized) neural net bayesian learning has a circuit prior — and thus, when learning a function which is implemented by some small circuit, only requires a small amount of training data to get good test accuracy — for certain scalings of the prior and with various other important caveats. The slides offer a self-contained presentation of the simplest version of the result. See the end of the presentation (slides 36–37) for a bunch of open problems in NN learning theory. Discuss
Score: 18🌐 MovesAug 19, 2026https://www.lesswrong.com/posts/SqDHeuycNkERurtSc/a-circuit-prior-in-nn-bayes - Yuxin Wu Examines Interpretable and Privacy-Preserving Methods for Trustworthy Artificial Intelligence
Yuxin Wu Examines Interpretable and Privacy-Preserving Methods for Trustworthy Artificial Intelligence USA Today
- Comu Launches Action Pro: An AI-Powered Hardware Device for Professional Meeting Workflows
Comu Launches Action Pro: An AI-Powered Hardware Device for Professional Meeting Workflows USA Today
- DocStyle Announces DocStyle AI: Document Productivity Tools, Accessible from the AI Assistants Lawyers Already Use
DocStyle Announces DocStyle AI: Document Productivity Tools, Accessible from the AI Assistants Lawyers Already Use USA Today
- Magneto IT Solutions Builds New AI Capabilities for Adobe Commerce Experiences
Magneto IT Solutions Builds New AI Capabilities for Adobe Commerce Experiences USA Today
- Enterprise AI platform Zenalyst raises pre-seed funding
Zenalyst, an agentic AI platform for enterprise execution, has raised Rs 3 crore (around $300K) in a pre-seed funding round from angel and corporate investors such as SKIL Cabs Private Limited, GERP Technologies Private Limited, Anurag Jain, Manish Kumar Jeloka, Pratap Padode, Siddarth Razdan and P H Corp. The fresh capital will be used to strengthen the ZenForce platform and expand its AI agent library across treasury, procurement and legal workflows, the Bengaluru-based company said in a press release. Founded in 2025 by Nagendra Singh, Sanketh Krishnappa and Vijay Jha, Zenalyst is building an AI platform for enterprise execution. Its platform, ZenForce, is designed to execute enterprise workflows across areas such as finance, procurement and legal operations. The company offers specialised AI agents such as ZenBank for enterprise money operations, ZenProcure for procurement and ZenLegal for contract intelligence. Zenalyst said its platform connects with more than 150 enterprise systems across ERP, CRM and banking. Zenalyst has onboarded clients such as Sattva Group, Blackstone invested Knowledge Realty Trust REIT, Bharat Biotech, Puravankara and Skill Travels. Its deployments are focused on sectors such as real estate, infrastructure, EPC, pharmaceuticals and travel. According to the company, its technology has reduced manual effort by up to 90% across targeted workflows, while select deployments have recorded payback periods of less than 12 months. The company currently has a sales pipeline of around $1 million. Zenalyst had earlier launched its AI Finance Workforce in October 2025 to target corporate finance functions. The platform was designed to connect data from systems such as ERP, CRM and HRMS and provide finance teams with automated analysis and decision support. In January 2026, the company said it had onboarded listed real estate developers as clients as it expanded its focus to asset-heavy sectors. Its product suite at the time included ZenBank, ZenBook, ZenForce and ZenPay, covering areas such as treasury, financial analysis, monitoring and payables. Zenalyst is headquartered in Bengaluru and has additional operations in Delaware and Chicago in the US. The company currently has a team of more than 40 people.
Score: 18💰 MoneyAug 19, 2026https://entrackr.com/snippets/enterprise-ai-platform-zenalyst-raises-pre-seed-funding-12394262 - How Manyavar Is Using AI To Take The Guesswork Out Of Retail
Manyavar is as much a technology and data company as it is an ethnic wear brand, claimed Vedant Modi, the…
Score: 18🌐 MovesAug 19, 2026https://inc42.com/buzz/how-manyavar-is-using-ai-to-take-the-guesswork-out-of-retail/ - 5 ways organisations are already winning with AI
5 ways organisations are already winning with AI uk.entrepreneur.com
Score: 18🌐 MovesAug 19, 2026https://uk.entrepreneur.com/technology/practical-ai-applications-for-organisations - How AI Helped The Bear House Grow Its Conversions 5X
As AI moves from powering search and recommendations to making decisions on behalf of consumers, ecommerce is beginning to shift…
Score: 18🌐 MovesAug 19, 2026https://inc42.com/buzz/how-ai-helped-the-bear-house-grow-its-conversions-5x/ - Avec’s new AI feature makes sure you never miss a deadline again
Avec's new Never Forget Anything feature automatically detects email deadlines and reminds you right when they matter, no manual setup needed.
Score: 16🌐 MovesAug 19, 2026https://www.digitaltrends.com/computing/avecs-new-ai-feature-makes-sure-you-never-miss-an-email-deadline-again/ - I Tried a Window-Cleaning Robot: Do Not Recommend
I tested the top-of-the-line Ecovacs Winbot W2S Omni window-cleaning robot on my Victorian home, and it was an unmitigated disaster.
- How Gwyneth Paltrow went from wellness tastemaker to AI acolyte
How Gwyneth Paltrow went from wellness tastemaker to AI acolyte Business Insider
Score: 15🌐 MovesAug 19, 2026https://www.businessinsider.com/gwyneth-paltrow-ai-era-tech-goop-new-obsession-2026-8 - Ex-Google chief scientist tells Gen Z the trick is not to master AI—instead, just ‘skim papers’
Ex-Google chief scientist tells Gen Z the trick is not to master AI—instead, just ‘skim papers’ Fortune
- A Boss Level Deal: Save $600 With This AI-Powered 18-Inch Acer Gaming Laptop
A Boss Level Deal: Save $600 With This AI-Powered 18-Inch Acer Gaming Laptop PCMag
Score: 15🌐 MovesAug 19, 2026https://www.pcmag.com/deals/boss-level-deal-save-600-18-inch-acer-gaming-laptop-august-18 - Public Sector Tech Report 2026: Local Government
How local governments are turning AI into advantage.
- An LLM wiki changed how I work
And everything else I learned about productivity this year
- Google’s rolling out an AI-themed visual tweak for paid accounts
Google's new colors are blue and purple.
Score: 15🌐 MovesAug 19, 2026https://www.androidauthority.com/google-account-profile-picture-3700664/ - Why founders should never use AI-generated pitch decks
I start my work weeks with deck reviews. Every Monday morning, in advance of our weekly fund meeting, I look at all the decks that entrepreneurs submitted through our website in the past week. Every deck gets a rating: “Yes!” means we should definitely take a meeting with the founders. “Maybe” means we should discuss first. “Too early” means we like the team and should watch their progress. “No” means, well, no. I am not a career VC. I was a founder and operator (read: held different jobs in tech) long before I ever made an investment. My journey to VC was heavily driven by a desire to open the heavy gates of venture to outsiders, misfits and mold-breakers like me. From the get-go, I promised myself I would give every entrepreneur a chance, whether they came highly recommended by an insider, or they came in cold through the “ Pitch Us ” link on the website. Which is why, every single Monday for the last five years, I’ve started my week looking at decks. We get anywhere from a dozen to a hundred in a week. My colleague David and I look at every single one. Our goal is not to learn everything about the startup in one sitting—our only goal is to decide whether we should meet them. A sea of sameness Inevitably, after the first few decks, it gets monotonous. It’s like grading papers— everything starts blurring together, even as I try very hard to focus and honor every deck with my attention. I wish this weren’t the case, but the fact is, the quality of these decks and the opportunity they are pitching varies broadly. The honest truth is that this part of the job can be a bit of a slog. Lately, it’s not a slog. It’s downright unbearable. You think LinkedIn is bad these days? Try sifting through three dozen AI -made pitch decks to start your week . A sea of sameness packed with buzzwords that say nothing. How do I know they’re made by AI? For one thing, they’re visually perfect—but that’s not the problem. The problem is they all have the same slides and are full of the same literary mannerisms you see on social media. And the slide titles: Always fragments. Never sentences. (See what I did there?) To be clear: I applaud entrepreneurs who use AI for things like research, to automate grunt work, polish their materials, and scale themselves. I do take issue with those who use AI to do the thinking and communicating for them. As an early-stage VC, an opportunity is only as special as the founder or founding team that’s leading it. If that founder can’t think for themselves … at a minimum, that’s just not very special at all. It’s the work, not the product At this point, you might argue, “Of course I’m thinking for myself! These are all my ideas! I just put them in prompts because VCs expect a perfect pitch deck and I don’t have time to make it perfect.” To which I would say: Whining gets you nowhere. First of all, while fancy graphics are nice, standout decks don’t need to be fancy—they need to be effective. McKinsey famously uses plain sans-serif font and a strict, minimalist approach to color for all their client work. When it comes to graphic design, a deck needs to be clean, not professionally designed. Not to mention—if you really hate making decks, you could write a short memo instead. Zero graphic design needed. Speaking of McKinsey: I worked there for a hot minute earlier in my career, followed by a stint at a boutique strategy consulting firm founded by Clay Christensen . It is no exaggeration to say that I spent that entire chunk of my life creating and perfecting decks. Every new client project was three months of “996” working on a single deliverable, inevitably a deck. We would agonize over every message and every image on every slide, and order and reorder those slides ad nauseum. Those decks were the bane of my existence. I had nightmares about them—still do! For a long time I questioned what was the point of the damn thing. I could hardly believe that clients were paying us so much money to make a friggin’ PowerPoint. Of course, they weren’t paying us for a deck. They were paying us to analyze a business problem and find the best possible solution. And we did that by building a storyline, breaking down the client’s problem into all its discrete pieces, testing hypotheses one by one, and iterating on the message until it was communicated perfectly. The final deck was a communication device. The work that went into it was what mattered. The same thing goes for entrepreneurs. A great pitch deck shows me you’ve researched, tested and thought deeply about every part of this opportunity, from the market to the customer to the marketing to the raise. It tells me you’re the expert, that you’ve done the work. As your prospective investor, why should I accept any less? How to cut through the noise Beyond the entrepreneurs’ basic handling of what it is they’re building, there’s an even bigger reason why a pitch deck is important: It shows that you know how to stand out from the crowd. Have you noticed how noisy the world is these days? There are ads and content literally everywhere. Everyone’s inboxes are packed with cold inbounds. Every possible marketing channel is saturated with brands. It is harder than ever to capture people’s attention. Not just investors—employees and customers, too. Every single startup—in fact, every single business—needs to figure out a way to stand apart. And it’s not just attention, it’s intention. Your entire job as an early stage startup founder is to convince people to make irrational decisions. You need to convince investors to give you money when there’s very little proof that the business is viable; you need to convince the most talented people to forego better salaries elsewhere and work for you; you need to convince customers to pay you for something that doesn’t really exist yet, nobody’s ever heard of or few have ever used. Do you honestly think you can do that with an AI-generated pitch deck? I want to say this as kindly as possible, but in the words of Brene Brown, clear is kind . So let me be clear: with the technology that’s available today, when you build your deck with AI, you sound exactly like everyone else . A problem of our own making Perhaps the proliferation of AI-slop pitch decks is a problem of the VC establishments’ own making, a logical consequence of decades of coaching founders about the importance of the perfect deck. Thousands of accelerator cohorts—including some of mine—have taught founders to obsess over a perfectly calculated TAM, SAM, and SOM, how to show your traction, and how to make a team slide that pops. And then as investors we get those decks that founders have spent countless hours on, we flip through in 20 seconds, and 99 times out of 100, we pass. No wonder founders are over it. I totally get their frustration. The flipside, of course, is that building a company is just plain hard. There are more founders than money, there’s more ideas than customers. Most of us don’t get to skip the hard part and sail straight to success. If you’re not willing to put in the work, you probably shouldn’t be doing this in the first place. Still with me? Let me tell you a secret I’ve been coaching founders on how to raise early-stage capital for a decade. In the very early days, I taught pitch decks like a checklist, just like everyone else did at the time. It wasn’t long before I realized that was entirely the wrong approach. The goal of a pitch deck is not to be comprehensive, it’s to be compelling. You don’t need to tell the investor everything up front, you need to share just enough to get a meeting. If you create enough excitement and curiosity through slides (or a memo, or even just a really great email), I promise you, the investor will want to ask about the rest. Mission accomplished. Your deck doesn’t need to be perfect. It doesn’t even need to be a deck! All it needs to do is answer three questions—why this, why you, and why now—as quickly, clearly and compellingly as possible. It is far, far harder to get this right than it is to write a prompt or even fill out a Canva template. I hope you do it anyway.
- NurseVeda integrates AI into nursing education platform
NurseVeda integrates AI into nursing education platform USA Today
Score: 15🌐 MovesAug 19, 2026https://www.usatoday.com/press-release/story/38994/nurseveda-integrates-ai-into-nursing-education-platform/ - The AI boom made San Francisco so crowded ‘tech bros’ making six figures are scrounging for homes
The AI boom made San Francisco so crowded ‘tech bros’ making six figures are scrounging for homes Fortune
Score: 15🌐 MovesAug 19, 2026https://fortune.com/article/ai-boom-san-francisco-tech-bros-six-figures-housing-08-10-2026/ - Synup debuts Sydekick AI: The Agent For Local Marketing
Synup debuts Sydekick AI: The Agent For Local Marketing USA Today
Score: 15🌐 MovesAug 19, 2026https://www.usatoday.com/press-release/story/40617/synup-debuts-sydekick-ai-the-agent-for-local-marketing/ - MoneyBuddy Launches Upgraded AI Loan-Matching Feature for SME, Personal and Mortgage Loans
MoneyBuddy Launches Upgraded AI Loan-Matching Feature for SME, Personal and Mortgage Loans USA Today
- Thais rush to register for TH-AI Passport
The Digital Economy and Society (DES) Ministry hopes the registration for the government's TH-AI Passport project will reach the target of 5 million before the project's launch on Aug 31.
Score: 15🌐 MovesAug 19, 2026https://www.bangkokpost.com/business/general/3304810/thais-rush-to-register-for-thai-passport - Save Up to $560 Off Mammotion Robot Lawn Mowers With App Control and AI Vision
Kick that old lawn mower to the curb and let your new robot take a load off.
- AuriQ Systems Launches Patent Analysis AI with Free Tier for Inventors
AuriQ Systems Launches Patent Analysis AI with Free Tier for Inventors azcentral.com and The Arizona Republic
- Hi3D Introduces AI-Powered Multi-Color Printing Workflow, Turning Digital Models Into Printable Creations
Hi3D Introduces AI-Powered Multi-Color Printing Workflow, Turning Digital Models Into Printable Creations USA Today
- Building a Decision Tree From Scratch — Understanding the Internal Working of it
This is the third entry in my “ML from scratch” series, after KNN and Gaussian Naive Bayes. This time: a decision tree classifier, built with nothing but NumPy, broken down from the concept all the way to individual lines of code. What Is a Decision Tree? A decision tree is a model that makes predictions by asking a sequence of yes/no questions about the input’s features. Each question narrows down the possibilities until you arrive at an answer. Structurally, it’s a binary tree: Every internal node holds a question of the form “is feature X ≤ some threshold?” Every leaf node holds a predicted class label. To classify a new sample, you start at the root, answer the question at each node, follow the corresponding branch (left if true, right if false), and repeat until you land on a leaf. Whatever label that leaf holds is the prediction. That’s it structurally. The interesting part — and the part that actually needs an algorithm — is how the tree decides which questions to ask, and in what order. Simple flow of how decision tree executes How It Works: The Core Idea Training a decision tree means building this tree of questions from data, one node at a time, top-down. At the root, you have your entire training set, with a mix of classes. The goal is to find one question — one (feature, threshold) pair — that splits the data into two groups that are, as much as possible, more homogeneous than the group you started with. Ideally, one side ends up mostly class A, the other mostly class B. Once you’ve picked that first question and split the data, you don’t stop — you treat each of the two resulting groups as its own smaller problem, and repeat the same process on each: find the best question to split that group further. This continues recursively, each split producing two more (smaller) groups to potentially split again. The recursion stops when further splitting stops making sense — a group is already pure (all one class), you’ve split enough times already, or there aren’t enough samples left to justify splitting further. At that point, the group becomes a leaf, and the leaf’s prediction is simply the majority class within it. So the whole algorithm is really just two ideas layered together: A way to score how good a candidate split is (so you can pick the best one at each node). Recursion — apply that scoring process over and over, on smaller and smaller subsets, until you’re out of useful splits to make. Everything else is implementation detail. The next section covers idea #1 — how “good” gets defined mathematically. Methodology: Entropy and Information Gain To score a split, we need a way to measure how mixed up (impure) a set of labels is — and how much a candidate split reduces that impurity. Decision trees commonly use entropy for the first part and information gain for the second. Entropy Entropy measures the disorder in a set of labels: E(S) = -Σ p(x) · log(p(x)) Where p(x) is the proportion of class x in set S, summed over all classes present. Worked example. A node with 10 samples: 6 of class 0, 4 of class 1. p(0) = 6/10 = 0.6 p(1) = 4/10 = 0.4 E = -(0.6 · log(0.6) + 0.4 · log(0.4)) = -(0.6 · (-0.511) + 0.4 · (-0.916)) = -(-0.3065 + -0.3665) = 0.673 Compare that to a pure node — 10 samples, all class 0: p(0) = 1.0 E = -(1.0 · log(1.0)) = -(1.0 · 0) = 0 Entropy of 0 means no disorder at all — nothing left to gain from splitting further. Entropy is at its maximum when classes are perfectly balanced, and shrinks toward 0 as one class comes to dominate the set. In code: def _entropy(self, y): hist = np.bincount(y) ps = hist / len(y) return -np.sum([p * np.log(p) for p in ps if p > 0]) Information Gain Entropy alone tells you how mixed a single set is. To evaluate a split , you compare the parent’s entropy against a weighted average of the two children’s entropy: IG = E(parent) - [ (n_left/n) · E(left) + (n_right/n) · E(right) ] Weighting by n_left/n and n_right/n matters: a split that carves off a large, pure group and leaves a small, mixed remainder should score differently than one that produces two medium, moderately-mixed groups. Weighting by size accounts for that. Worked example , continuing the 10-sample node (parent entropy E = 0.673). Suppose a candidate split sends 5 samples left (all class 0) and 5 right (1 class-0, 4 class-1): E(left) = 0 (pure) E(right) = -(0.2·log(0.2) + 0.8·log(0.8)) = -(0.2·(-1.609) + 0.8·(-0.223)) = -(-0.322 + -0.179) = 0.500 Weighted child entropy = (5/10)·0 + (5/10)·0.500 = 0.250 IG = 0.673 - 0.250 = 0.423 That’s a strong split — it fully isolated a pure group. A weak split, where the class ratio barely changes on either side, produces an IG close to 0. In code: def _information_gain(self, y, X_column, threshold): parent_entropy = self._entropy(y) left_idxs, right_idxs = self._split(X_column, threshold) if len(left_idxs) == 0 or len(right_idxs) == 0: return 0 n = len(y) n_l, n_r = len(left_idxs), len(right_idxs) e_l, e_r = self._entropy(y[left_idxs]), self._entropy(y[right_idxs]) child_entropy = (n_l/n) * e_l + (n_r/n) * e_r return parent_entropy - child_entropy Choosing the Best Split With a way to score any single candidate split, finding the best split at a node is a search: try every feature, and for each feature, try every value that appears in that feature’s column as a candidate threshold, keeping whichever (feature, threshold) pair produced the highest information gain. best_split(X, y) = argmax over all (feature, threshold) of IG(y, X[:, feature], threshold) Only unique observed values need to be tried as thresholds — any value strictly between two consecutive observed values produces an identical split, so there’s no benefit to checking a finer-grained range. Walking Through the Code, Part by Part Node class Node: def __init__(self, feature=None, threshold=None, left=None, right=None, *, value=None): self.feature = feature self.threshold = threshold self.right = right self.left = left self.value = value def is_leaf_node(self): return self.value is not None Node is deliberately dual-purpose — the same class represents both internal nodes and leaves, distinguished only by which fields are set: An internal node has feature and threshold (the question it asks) plus left and right (its two children). value stays None. A leaf node has only value set (the predicted class). Everything else stays None. value is keyword-only (the * forces this) specifically so it can't be passed positionally by accident and confused with left/right — leaves and internal nodes are constructed with visually distinct calls: Node(value=leaf_value) vs. Node(best_feature, best_threshold, left, right). is_leaf_node() checks self.value is not None — which works because internal nodes never set value, and leaves never set anything else. This one check is what predict uses to decide whether to stop traversing or keep going. __init__ def __init__(self, min_samples_split=2, max_depth=100, n_features=None): self.min_samples_split = min_samples_split self.max_depth = max_depth self.n_features = n_features self.root = None Three hyperparameters, all controlling when the tree stops growing (directly or indirectly), plus self.root, which starts empty and gets filled in by fit. min_samples_split — the minimum number of samples a node needs to be eligible for splitting at all. max_depth — a hard cap on how many splits deep the tree can go. n_features — how many features to consider at each split. None means "use all of them"; a smaller number introduces the random feature subsampling used by Random Forests. fit def fit(self, X, y): self.n_features = X.shape[1] if not self.n_features else min(X.shape[1], self.n_features) self.root = self.grow_tree(X, y) The entry point. The first line resolves n_features into an actual number: if none was specified, use every feature in X; if one was specified, use whichever is smaller — the requested count or the number of features actually available (so you can't accidentally ask for more features than exist). The second line kicks off recursive tree-building and stores the resulting root node. grow_tree def grow_tree(self, X, y, depth=0): n_samples, n_feats = X.shape n_labels = len(np.unique(y)) if depth >= self.max_depth or n_labels == 1 or n_samples < self.min_samples_split: leaf_value = self.most_common_label(y) return Node(value=leaf_value) feat_idxs = np.random.choice(n_feats, self.n_features, replace=False) best_feature, best_threshold = self.best_split(X, y, feat_idxs) left_idxs, right_idxs = self._split(X[:, best_feature], best_threshold) left = self.grow_tree(X[left_idxs, :], y[left_idxs], depth + 1) right = self.grow_tree(X[right_idxs, :], y[right_idxs], depth + 1) return Node(best_feature, best_threshold, left, right) Each call handles one node, given the slice of data (X, y) that reached it and how deep it is (depth). Stopping check first. The if condition covers the three ways a node becomes a leaf: it's too deep (depth >= max_depth), it's already pure (n_labels == 1), or it has too few samples to keep splitting (n_samples < min_samples_split). If any is true, skip straight to most_common_label(y) and return a leaf — no point searching for a split that won't be used. Otherwise, split and recurse. np.random.choice(n_feats, self.n_features, replace=False) picks which features this node is even allowed to consider — all of them by default, a random subset if n_features was restricted. best_split searches those features for the best (feature, threshold) pair. _split then partitions the actual data into left_idxs/right_idxs based on that choice. The two recursive calls — self.grow_tree(X[left_idxs, :], y[left_idxs], depth + 1) and the equivalent for right — are where the "repeat the whole process on each smaller group" idea from earlier actually happens in code. Each call returns a fully built subtree (root node of that subtree), and the final line wires both of those subtrees into a new internal Node, which is what gets returned up to this call's caller. That's how depth-first construction bubbles all the way back up to a single root. best_split def best_split(self, X, y, feat_idxs): best_gain = -1 split_idx, split_threshold = None, None for feat_idx in feat_idxs: X_column = X[:, feat_idx] thresholds = np.unique(X_column) for thr in thresholds: gain = self._information_gain(y, X_column, thr) if gain > best_gain: best_gain = gain split_idx = feat_idx split_threshold = thr return split_idx, split_threshold A brute-force search, directly implementing the argmax from the methodology section. The outer loop goes feature by feature; the inner loop goes threshold by threshold (every unique value observed in that feature's column). For every (feature, threshold) pair, _information_gain scores it, and best_gain/split_idx/split_threshold track the best one seen so far. best_gain starts at -1 specifically because information gain is always ≥ 0 — so the very first real split evaluated is guaranteed to beat the initial placeholder and get recorded. _information_gain and _split def _information_gain(self, y, X_column, threshold): parent_entropy = self._entropy(y) left_idxs, right_idxs = self._split(X_column, threshold) def _split(self, X_column, split_thresh): left_idxs = np.argwhere(X_column <= split_thresh).flatten() right_idxs = np.argwhere(X_column > split_thresh).flatten() return left_idxs, right_idxs These two are the direct code form of the entropy/information-gain formulas covered above — _split produces the two index arrays for "at or below the threshold" vs. "above it" using np.argwhere and a boolean comparison; _information_gain uses those indices to compute each side's entropy and combine them into the weighted score. The one piece of defensive logic — if len(left_idxs) == 0 or len(right_idxs) == 0: return 0 — handles a threshold that doesn't actually separate anything (every sample lands on one side), which would otherwise divide by an empty array when computing that side's entropy. _entropy and most_common_label def _entropy(self, y): hist = np.bincount(y) ps = hist / len(y) return -np.sum([p * np.log(p) for p in ps if p > 0]) def most_common_label(self, y): counter = Counter(y) return counter.most_common(1)[0][0] _entropy is the formula from the methodology section, translated line for line: np.bincount(y) counts samples per class, dividing by len(y) turns those into proportions, and the list comprehension sums p · log(p) over every class with p > 0 (skipping zero-count classes, since log(0) is undefined and they contribute nothing anyway). most_common_label is what a leaf actually predicts: Counter(y).most_common(1) returns the single most frequent label in y as a (label, count) tuple, and [0][0] pulls out just the label. predict and _traverse_tree def predict(self, X): return np.array([self._traverse_tree(x, self.root) for x in X]) def _traverse_tree(self, x, node): if node.is_leaf_node(): return node.value if x[node.feature] <= node.threshold: return self._traverse_tree(x, node.left) return self._traverse_tree(x, node.right) predict just runs _traverse_tree on every row of X and collects the results into an array. _traverse_tree is the inference-time mirror of the "start at the root, answer the question, follow the branch" description from the intro: check if the current node is a leaf (if so, return its stored label — done); otherwise, compare x[node.feature] against node.threshold and recurse into node.left or node.right accordingly. No computation happens here — training already decided every question in advance, so prediction is pure traversal, taking O(depth) steps per sample regardless of how large the training set was. Results Trained and evaluated on sklearn.datasets.load_breast_cancer (80/20 split, random_state=1234, default hyperparameters): Accuracy: 94% That’s competitive with sklearn’s own DecisionTreeClassifier on the same split — expected, since the underlying math is identical. sklearn's version is faster thanks to more optimized split-finding, and supports additional criteria (Gini impurity) and pruning options this implementation doesn't. To go beyond the raw accuracy number, I added a few visualizations: a confusion matrix to see where the errors landed, a feature-importance chart based on how often each feature was used to split (which features the tree actually relied on), and a PCA projection of the test set colored by correct vs. incorrect predictions, to check whether misclassifications clustered in any particular region of the data. What This Implementation Doesn’t Do No pruning. A fully-grown tree (especially with max_depth=100) can overfit badly on noisier data. Real implementations add pre-pruning (stricter stopping criteria) or post-pruning (cost-complexity pruning, ccp_alpha in sklearn). No Gini impurity option. Entropy is one valid splitting criterion; Gini impurity is another, cheaper to compute (no logarithms) and commonly used as a default in practice. Exhaustive threshold search. Trying every unique value in every feature works fine for a learning implementation, but scales poorly — sklearn’s C-optimized version does the same search far more efficiently. No built-in feature importance. The version I computed for the plots is a simple split-frequency count, not the sample-weighted (samples × information gain) importance sklearn reports. Next Steps The natural extension from here is a Random Forest — since n_features subsampling is already built in, most of the ensemble machinery is really just: train many of these trees on bootstrapped samples of the data, and aggregate their predictions by majority vote. That's next in this series. Complete Code of Decision Tree implementation is available on my GitHub. GitHub link — https://github.com/Archan47/Machine-Learning-Algorithms-From-Scratch Code for this and the rest of the “ML from scratch” series is on GitHub. Building a Decision Tree From Scratch — Understanding the Internal Working of it was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.