AI News Archive: August 19, 2026 — Part 24
Sourced from 500+ daily AI sources, scored by relevance.
- localmd
An AI knowledge base in your browser and your local folder
- ZEON Format
Stop wasting money on JSON. Save 50%+ tokens on your LLMs
- CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks
Most LLM benchmarks rank models on their ability to automate work tasks. In practice, however, models are often used to assist other (human or LLM) agents. The question that drives model selection is therefore not only which model produces the best output, but which model most improves the work of a...
- Bayesian Partner Modelling enables Adaptive Replanning for LLM Coordination
Multi-agent Large Language Model (LLM) systems often struggle to collaborate with new teammates whose strategies shift mid-task. Because agents execute multi-step or temporally extended skills, they frequently continue executing outdated plans long after public evidence shows that a partner has chan...
- A Locally Deployable Tool-Grounded LLM Multi-agent Framework for Automating Methane Emission Analysis and Reporting
Methane field monitoring requires the integration of sampling design, meteorological interpretation, sensor processing, plume analysis, visualization, and reporting, but these steps are often distributed across separate expert-driven workflows. We developed a locally deployable, tool-grounded large ...
- Autonomous Cyber Defense in Connected Vehicles: A Multi-Agent Approach to V2X Security
A connected vehicle has roughly 100 milliseconds to decide whether an incoming Basic Safety Message is real or fabricated. If a false emergency braking alert reaches the planning pipeline in time, the car brakes - a safety failure triggered by a security failure. Existing intrusion detection systems...
- TractorBeam: Personalized AI Sensemaking Support via Collaborative Machine Annotation
Language model-based systems which allow asking questions of documents have become popular tools for sensemaking. Despite their implied capability, these systems still suffer from issues of factuality and provenance, while encouraging confirmatory, rather than exploratory, research. We present Tract...
- Zero-Shot SAM2 Segmentation and Vision Transformer-Based Recognition of Elamite Cuneiform Symbols from Degraded Tablet Images
Automated recognition of ancient cuneiform script poses a compound signal-degradation problem: the three-dimensional relief of clay tablets creates spatially varying illumination and cast shadows, surface erosion introduces structured noise that overlaps with genuine sign impressions, and severe cla...
- Reducing Technician Search Burden: A Multimodal RAG for Cessna 172 Maintenance Manual
Proper use of the aircraft maintenance manual is essential for correct maintenance, providing procedures, diagrams, cautions, and specifications. However, technicians often avoid consulting it because it is difficult to navigate and time-consuming under strict schedules. Retrieval augmented generati...
- LEDGER: Claim-to-Evidence Trace Graphs for Auditing LLM Agents
Large language model (LLM) agents can now carry out long-horizon technical workflows involving complex tool use, code execution, file edits, and generated artifacts. As agents do more work faster, the productivity bottleneck shifts from producing outputs to auditing whether those outputs are correct...
- LearnAI: Just-in-Time AI Co-Creation Across Disciplines at a University
As generative AI reshapes professional and educational practice, institutions face a challenge: how to support diverse learners, from non-coders to advanced students, in building confidence and practice with AI-supported problem solving. Most institutional responses bifurcate into conceptual worksho...
- Report on The 1st Workshop on Human-Centered Proactive and Personalized Agents for Interactive Information Access at CHIIR 2026
Interactive information access is increasingly moving beyond reactive query-response paradigms toward agentic systems that can personalize interaction, retain context, infer latent needs, recommend next steps, and initiate support. This shift creates new opportunities for adaptive and context-aware ...
- SemanticSlider3D: Training-Free Continuous Semantic Editing for 3D Objects
Fine-grained control over continuous semantic attributes of 3D objects is essential for 3D content creation, but is not well supported by conventional 3D modeling workflows or prompt-based interaction with existing generative AI tools. While slider-based methods have proven effective for fine-graine...
- SiNMULI: Novel Signed Network Approach for Malicious URL Identification
In today's era of rapid advancements in artificial intelligence, computer security and online safeguarding measures have undergone significant improvements. However, malicious websites continue to facilitate the spread of phishing schemes, fraudulent activities and unsolicited communications. Conven...
- Malformer: A Multi-Modal Malware Detector Using Transformers
Traditional malware detection systems that rely on a single representation of malware often fail to identify novel threats. These representations of malware binaries, also known as modalities, do not provide the models with sufficient information to discriminate among all samples. Additionally, indi...
- Catastrophic Learning: A New Attack Vector on Continual Learning Networks
Continual Learning (CL) enables deep learning models to iteratively learn from a stream of data without forgetting prior knowledge. Existing adversarial research on CL primarily aims to re-enable catastrophic forgetting, attacking stability and reducing availability. We identify a novel security fla...
- FedLNS: Leverage LayerNorm Signature Modeling to Mitigate Adversarial Manipulation in Federated LLMs
Federated training enables language models to learn from distributed private text, but the server cannot directly verify the local supervision or optimization process that produces each client update. A malicious client can therefore train on corrupted targets, introduce incorrect context-token asso...
- Improving LLM-Based SSH Honeypots Through Prompting and Fine-Tuning
LLM-based SSH honeypots often use closed cloud LLMs because they give strong shell realism, but cloud models create deployment problems. These include no stable versioning, provider-side changes, attacker-driven cost, and model decommissioning. Local open-weight models avoid these problems, but they...
- CTIFoundry: An Agent-Native Corpus Scaffold for Cyber Threat Intelligence
Cyber threat intelligence (CTI) is increasingly consumed not by human analysts but by LLM agents that compose multi-step investigations at query time. The harness side of this shift has matured rapidly (planning loops, tool protocols, context management), but the corpus side has not: threat reports ...
- Denoising-Aware Inversion: Revealing Privacy Risks in Noise-Protected Text Embeddings
Dense text embeddings are widely used in data mining, retrieval, and downstream machine learning systems due to their compact and semantically rich representations, but recent embedding inversion attacks have shown that they can expose substantial information about the original text, leading to seri...
- FedGuard-DC: Privacy-Preserving Federated Load Forecasting and Cyber-Attack Detection for Data-Center Loads in Transmission Systems
The rapid growth of large data-center (DC) loads is creating new challenges for power-system visibility, privacy, and cyber-physical security. System operators need accurate short-term information about these fast-varying loads, while DC operators may avoid sharing raw megawatt measurements because ...
- Toward Quantum Advantage in Learning Parities with Structured Noise via Lower Bound Optimization of the Condition Number
Learning Parities with Structured Noise (LPSN) can be reduced to solving nonlinear Boolean systems. In quantum computing, such systems are typically transformed into Macaulay linear systems and solved via quantum linear system algorithms, a process severely limited by the condition number. To addres...
- Who Can Make the Action Happen? An Authority-Decomposition Framework for High-Risk Automated Systems
High-risk automated systems distribute control across services, credentials, protected components, and lifecycle mechanisms. Labels such as authorized, approved, privileged, or protected therefore do not answer a basic causal question: which actors can actually make a consequential action occur? Thi...
- CauSec: Unboxing the Causal Drivers of Static Vulnerability Analysis Performance
Static Application Security Testing (SAST) tools are widely used in both industry and academia. Such tools often make design choices that sacrifice detection to achieve higher performance, i.e., increased precision, decreased runtime, or increased scalability. These design choices rely on certain as...
- Decisive Margins in Differentially Private Voting
Differential privacy protects individual voting records by injecting randomness into the published outcome, but this noise can lead to erroneous results when an election is close. We study how precise central differential privacy and local differential privacy can be for common voting rules, includi...
- Geometric Data Perturbation with Noisy-Anchor Alignment for Privacy-Preserving Collaborative Learning
Geometric Data Perturbation (GDP) enables one-shot, privacy-preserving collaborative learning: each participant applies a distance-preserving transformation to its private data and uploads only the resulting representation to a central analyst. We study GDP under analyst-participant collusion, in wh...
- IriSig-Spoof: A Real-World Benchmark for Time-Robust Satellite RF Fingerprinting and Spoofing Detection
Low Earth orbit (LEO) satellite Internet is becoming critical communications infrastructure, yet its open wireless links remain vulnerable to satellite impersonation and signal spoofing. Radio frequency fingerprinting (RFF) offers a potential defense by exploiting transmitter-specific hardware imper...
- VQC-ZTI: Variational Quantum Control for Zero Trust Protection of the Tactile Internet
Tactile Internet services couple cyber events directly to physical actuation, so security decisions must improve risk discrimination without perturbing the control path. This paper presents VQC-ZTI, a split-plane Variational Quantum Classifier framework for zero-trust protection of Tactile Internet ...
- Beyond Distortion Robustness: Rethinking Severe Cropping as Erasure-Resilient Message Embedding
Robust message embedding in images is important for multimedia security applications such as copyright protection and content tracing. Existing methods are largely developed under a distortion robustness paradigm, where the embedded signal remains spatially present but is degraded by noise, blur, or...
- Flama: a Python framework for development and deployment of production-ready APIs, machine learning, and LLM services
We present Flama, an open-source Python framework for developing and deploying production-ready web APIs, machine learning services, and large-language-model (LLM) applications. Built on the Asynchronous Server Gateway Interface (ASGI), Flama offers a type-driven, async-first programming model that ...
- Code Health in LLM-Based Test Generation: Effectiveness and Token Efficiency
Coding agents powered by Large Language Models (LLMs) are now prominent in software engineering. Previous work has shown that AI tools perform better on high-quality source code that is easy to maintain. In this study, we investigate how the effectiveness of LLM-generated unit tests varies across ma...
- OdinEval: A Reproducible Benchmark for LLM-Based Program Repair in the Odin Programming Language
Repository-level repair benchmarks still center on a few mainstream languages, leaving systems languages such as Odin largely untested. We present OdinEval, a reproducible benchmark built from documented defects in public Odin repositories. Each instance binds an issue to base and fix commits, a gol...
- AppEval: A Unified Benchmark for LLM-Based Mobile Application Repair in ArkTS, Swift, and Kotlin
Repository-level LLM agents are typically evaluated on projects whose tests run on the build host. It remains unclear whether their repairs survive the mobile build-install-launch-test boundary, where a missing SDK, offline device, or pre-assertion crash can be mistaken for a program failure. We pre...
- SemaPLC: A Project-Grounded, Verification-Gated Agent Harness for PLC Code Generation
Programmable logic controllers (PLCs) run industrial plants, and large language models can already generate independent program organization units (POUs) for them. Whether such logic integrates into an existing PLC project and then runs correctly has been checked only in limited tests. We present \t...
- Think-to-Personalize: Unifying Reasoning and Retrieval for User-Centric Personalized Dense Retrieval
Dense retrieval has become a cornerstone of modern local-lifestyle e-commerce search by encoding queries and items into semantic embedding spaces. While recent advancements have transitioned from BERT-based embedding models to Large Language Models (LLMs), most approaches still treat LLMs as static ...
- GateDiffInt: Gate-Mediated Controllable Diffusion and Multi-Intent LLM Distillation for User Behavior Modeling
Existing ranking models encode intent only implicitly, making it hard to disentangle structured intents of varying strength and temporal scale. Noise and intent in behavior sequences are mutually reinforcing---we call this Noise--Intent Coupling (NIC). Noise dilutes true intents, while the lack of s...
- Visual-Aware Representation of Web Pages for Machine Learning Applications
Applying machine learning to web pages is challenging due to the need to interpret HTML together with associated resources and perform rendering to obtain a meaningful visual and layout-aware representation. As a result, machine learning over web content remains comparatively underexplored. In this ...
- SIDScope: A Diagnostic Resource for Semantic-ID Interfaces in Generative Recommendation
Semantic-ID mappings are reusable interfaces between item tokenizers and generative recommenders, yet released mappings rarely state whether they are coherent, what structure they expose, how generated paths resolve, or what must be revalidated after a refresh. SIDScope is a source-traced diagnostic...
- OneModel: A Unified Foundation for Platform-Scale Multi-Scenario Ranking
Platform-scale recommender systems often span multiple business streams such as organic recommendation, advertising, and merchant services, where user behaviors form a continuous cross-stream trajectory. Maintaining separate ranking systems fragments user representations and increases engineering co...
- FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI Systems
Large language models are increasingly used to support financial operations, but their apparent reasoning performance can depend on whether they receive the right evidence. In financial reconciliation, the evidence needed for diagnosis is distributed across invoices, purchase orders, approvals, allo...
- More Context, Same Budget: Dual-Bounded Relational Recall Beyond Top-K Retrieval
More context does not require a larger retrieval budget. Under the same ceiling, a retrieval system can recover more of the evidence a question requires by following relationships between evidence that flat top-k ranking leaves behind. We test that proposition with Dual-Bounded Relational Recall (DB...
- Signature Recontextualization: Mapping perturbational signatures across biological contexts
Perturbational transcriptomics is a powerful tool for understanding gene function and drug effects, yet predicting how perturbations manifest across different biological contexts remains a central challenge, limiting translation from model systems to clinically relevant tissues. Despite growing interest in this problem, benchmarking efforts have been hindered by inconsistent evaluation tasks, heterogeneous metrics, and limited assessment across perturbation types and biological systems. Here, we introduce a benchmarking framework for cross-context perturbation-signature prediction (a task we define as signature recontextualization), grounded in explicit definitions of the prediction task, target-data availability, and evaluation metrics centered on signature recovery. The framework evaluates prediction performance across three target-context data regimes: (1) control only, where only control profiles from the target context are measured; (2) low coverage, where a limited subset of perturbations in the target context are measured; and (3) high coverage, where most perturbations in the target context are measured. This design enables systematic assessment of how prediction performance depends on target-context sample size while providing a standardized basis for comparing methods. We evaluate newly developed projection-based (projectCor) and network-based (netProp) methods alongside deep learning-based foundation models (scGPT, STACK) and statistical baselines. The benchmark spans four diverse perturbational datasets: CRISPR knockdowns and drug perturbations in cell lines, plus in vivo chemical perturbations in rat tissues from DrugMatrix, extending evaluation beyond isolated cell-line models to tissue-level responses. Across tasks, projection and network propagation approaches show strong flexibility across perturbation types and biological contexts, and in several cases match or exceed the performance of deep learning and foundation models, suggesting that model complexity does not inherently improve cross-context generalization. We further show that perturbation predictability varies substantially with pathway conservation, transcriptional response strength, and baseline similarity between source and target contexts. All datasets, methods, and evaluation utilities are released as an open-source R package (sigRecon), providing a foundation for reproducible benchmarking and future method development.
- OmegaSwitch: Bayesian Markov-Modulated Codon Models for Estimating dN/dS
Selective pressures can vary across both sites and evolutionary lineages; however, most codon models accommodate heterogeneity along only one of these dimensions and require the number of selective regimes to be specified in advance. Here, we introduce OmegaSwitch, a Bayesian phylogenetic software framework for inferring changes in the nonsynonymous-to-synonymous substitution-rate ratio (dN/dS) across sites and through evolutionary time. We implement a Markov-modulated codon model in which lineages transition among discrete dN/dS regimes and use reversible-jump Markov chain Monte Carlo to infer the number of regimes simultaneously. We further develop a Dirichlet-process mixture extension that allows the parameters governing these time-heterogeneous processes to vary among sites. Ancestral sampling produces joint posterior distributions of dN/dS across sites and nodes of the phylogeny, enabling lineage- and site-specific summaries with quantified uncertainty. Simulation analyses showed that both the posterior intervals for dN/dS and the number of evolutionary regimes were well calibrated under both models. We demonstrate OmegaSwitch using vertebrate alpha- and beta-globins. OmegaSwitch therefore provides a flexible Bayesian framework for investigating how selective pressures vary across protein-coding sequences and phylogenetic history.
- Brain-Language Alignment During Naturalistic Reading and Its Disruption by Mind-Wandering
Encoding models offer a principled framework for linking computational representations of language to neural activity, but most electroencephalography (EEG) evidence for brain--language alignment comes from tightly controlled, word-by-word reading paradigms. Whether such alignment is detectable during naturalistic reading, and how it is affected by lapses in attention, remains unclear. We addressed these questions using ROAMM, a multimodal dataset containing simultaneous EEG and eye-tracking recordings with time-resolved mind-wandering (MW) annotations from 44 participants reading naturalistic texts. Ridge regression encoding models were trained to predict fixation-aligned EEG spectral power and fixation-related potentials (FRPs) from five word-embedding models (GloVe, word2vec, BERT, GPT-2, and Llama 3). Using permutation testing with false discovery rate correction, we found statistically reliable brain--language alignment across both feature types, with contextual embeddings outperforming static embeddings. Spectral alignment was strongest in the alpha and low-beta bands over parietal electrodes, while FRP-based alignment peaked 200--300 ms after fixation onset over central and parietal-occipital regions. Leveraging ROAMM's span-level MW annotations, we further show that brain--language alignment is systematically reduced during MW, an effect that was substantially larger for oscillatory (PSD) than for event-related (FRP) features. These findings demonstrate that modern language-model representations are reflected in EEG activity during naturalistic reading despite the modality's inherent noise, and that fluctuations in attention constitute an underappreciated source of variability in brain--language encoding studies.
- Transmembrane coupling of protein condensates via membrane-mediated interactions: A simulation study
Recent experiments show that protein condensates sitting on opposite surfaces of a flat lipid membrane move together and prefer to overlap, even though they cannot touch each other. This points to an indirect, membrane-mediated interaction. Two mechanisms could be responsible: a curvature-induced interaction, which is energetic in origin, and a fluctuation-induced interaction, which is entropic. Here we study both with coarse-grained molecular dynamics simulations, using Cooke's implicit-solvent lipid model together with a generic bead-spring polymer model for the condensate. We compute the potential of mean force between two condensates across the membrane. For condensates of the same size, full overlap is unfavorable, and the pair instead settles into a partially overlapping state that bends the membrane into an S-like shape. When the two condensates differ strongly in size, full overlap becomes favorable. We explain this with a simple geometric picture. The condensate wets the membrane as a thin film and imposes curvature only along its rim, while membrane tension flattens the membrane under its interior. The resulting ring of curvature can trap a smaller condensate on the opposite side. We also compare the bending undulations and the effective bending modulus of a bare membrane, a membrane with one condensate, and a membrane with condensates on both sides. A wetting condensate suppresses the undulation modes and stiffens the membrane, but whether this makes overlap entropically favorable remains inconclusive. Our results indicate that the coupling is driven mainly by curvature, and that it depends on the wetting mechanism and on the membrane tension.
- ImpRes: A robust FRAP framework to quantify fast diffusion of cytoplasmic probes
Diffusion within the cytoplasm is fundamental to numerous biological processes. Fluorescence recovery after photobleaching (FRAP) is one of the most common method for quantifying molecular diffusivity in living cells using standard laser scanning confocal microscopy (LSCM). However, accurately measuring fast cytoplasmic diffusion (typically >10 m^2/s) is challenging due to rapid recovery kinetics, weak signal-to-noise ratios, post-bleach signal artifacts, and spatial restrictions affecting normalization. While individual challenges have been addressed in specific contexts, a simple and robust framework to quantify cytoplasmic diffusivity remains elusive. Here, we present a FRAP methodology specifically designed to overcome these obstacles. By utilizing the Gaussian function -- the impulse response (ImpRes) of the diffusion equation in an infinite medium -- our approach leverages the full spatiotemporal dataset through a single-equation three-parameter fitting procedure, thus releasing restrictions to small regions of interest and arbitrary initial time-points. The methodology was validated on three datasets of increasing complexity: in silico simulated recovery profiles, in vitro data from FITC-dextran in glycerol solution, and live-cell imaging of free cytoplasmic GFP. Systematic comparison with existing models demonstrates that the ImpRes approach significantly reduces sensitivity to noise and imperfect fluorescence normalization, while remaining robust against short-term biases, such as transient probe photo-activation. Given its robustness under realistic experimental conditions and its ease of implementation, the proposed FRAP methodology provides a reliable tool for quantitative cytoplasmic analysis.
- An Exact-Residue Atlas of Opioid Receptor Wiring and Rewiring across Ligand and Transducer Contexts
Opioid-receptor structures span four human receptor subtypes, diverse ligands, signaling partners, and experimental constructs. We curated 86 human opioid-receptor structures representing 84 independent experimental maps, with one unique experimental data set counted once for structure-level inference, and analyzed them using our in-house StrucMind platform. StrucMind constructs exact Ballesteros-Weinstein (BW) contact graphs, meaning residue-contact networks restricted to unambiguous generic BW positions. Relative to active transducer-bound structures, structures classified as inactive showed 2.49% lower mean contact similarity and 34.84% more rewired contacts, where rewiring is the static set of contacts gained or lost between two structures. Among 77 maps with a resolved selected-ligand site, changed contacts were 15.28% direct to the site, 40.13% adjacent at one graph edge, and 44.59% connected-distal at a finite graph distance greater than one. The deposited-water analysis identified 137 receptor-proximal waters. Sixty-one contacted at least two protein residues, including 38 that bridged at least two exact-BW residues; a separate ligand-contact branch contained 10 waters contacting both selected ligand and receptor, only 3 of which belonged to the 38-water set. None of 32 component-association tests survived global correction. For peptide versus small molecule, the smallest nominal p value among four outcomes corresponded to 6.18% lower shared-contact distance root-mean-square deviation (p=0.00989; q=0.3165, where q is the adjusted p value). The atlas supports bounded, testable hypotheses, not causal component, hydration, or efficacy mechanisms.
- AlphaVaR: an R framework for the statistical interpretation of AlphaGenome variant-effect predictions
AlphaGenome (Google DeepMind) scores a DNA variant across thousands of functional tracks at single-base resolution, reporting both the magnitude of each predicted effect and its rarity against a genome-wide background. That volume is itself the obstacle to biological interpretation. Here we present AlphaVaR, an R package that gives AlphaGenome's output a typed structure together with the statistical methods and visualizations needed to interpret it. The output schema is identical for every variant, so the same tests apply throughout it. AlphaVaR provides localization tests with multiple-testing correction and effect sizes, a specificity index measuring how far an effect concentrates on a few elements of a chosen variable, and a transparent prioritization that ranks candidates across interpretable criteria and maps each to a target gene. Results feed a plot library, reproducible reports and a code-free Shiny application. Applied to rs1427407, the lead common variant for fetal-haemoglobin level, AlphaVaR recovers the established biology of the BCL11A erythroid enhancer. Availability and implementation: https://github.com/KarimMarhaba/AlphaVaR, released under the MIT licence, R [≥] 4.2, with documentation at https://karimmarhaba.github.io/AlphaVaR/. The released version is archived at Zenodo (doi:10.5281/zenodo.21939265); the AlphaGenome scores analysed here are archived as a separate dataset (doi:10.5281/zenodo.21920988), and the scripts that regenerate every figure and reported number are in the repository (Supplementary Section S5).
- AlphaGenome deletion responses complement supervised enhancer-gene relation prediction in primary human astrocytes
Sequence-to-function models predict molecular readouts directly from DNA, but recognizing a functional regulatory element is not equivalent to assigning the gene it regulates. We evaluated whether AlphaGenome deletion responses identify experimentally supported enhancer-gene relations, using a frozen K562 analysis and a primary-human-astrocyte CRISPR interference (CRISPRi) resource external to our analysis. The K562 mean contrast was positive but heavy-tailed, and exact joins showed direct collision with released Gasperini and ENCODE-rE2G resources; we therefore treated K562 as supporting evidence. In a frozen evaluation of 2,307 AstroREG relations, AlphaGenome deletion strength discriminated 133 functional relations from 2,174 well-powered nonfunctional relations (average precision 0.479, enhancer-cluster 95% confidence interval 0.394-0.561, prevalence 0.058; area under the receiver-operating-characteristic curve 0.726, 0.659-0.786). Adding AlphaGenome to distance, ABC score, enhancer length, measured expression and assay-depth context increased enhancer-grouped out-of-fold average precision from 0.396 to 0.534 and improved log loss from 0.169 to 0.150. The authors' cross-fitted EGrf score was stronger alone (average precision 0.559); in a post-hoc calibration that held out both gene and enhancer folds, adding AlphaGenome increased average precision from 0.550 to 0.619 (paired enhancer-cluster increment 0.068, interval 0.023-0.115) and improved log loss from 0.143 to 0.132. This comparison had asymmetric inputs: EGrf was supervised on AstroREG labels and used local epigenomic and context features, whereas the AlphaGenome score was not fitted in this study to those labels or that feature panel but was read from a pre-existing primary-astrocyte RNA-seq output track. A post-hoc same-enhancer analysis gave conditional AUC 0.741 (0.663-0.814); a smaller same-gene analysis (34 genes, 155 relations) gave 0.701 (0.571-0.823). AstroREG labels and EGrf outputs were public before AlphaGenome's public release, so this evaluation is external to our study but not a post-release or proven-unseen benchmark. The results support complementary relation-level utility, not EGrf superiority, sequence-only deployment, causal assignment at arbitrary loci or equivalence between sequence deletion and CRISPRi.
- Use of the Elston-Stewart algorithm for the efficient calculation of exact pedigree-based Y-STR match probabilities
The formal assessment of a genetic match between a suspect and some biological trace material is one of the key tasks of forensic genetics, particularly in cases of sexual offence. The analysis of Y-chromosomal short tandem repeats (Y-STRs) has proven especially useful in this context. For a long time, however, calculating the probability of a perfect Y-STR profile match under the defense hypothesis that the suspect was not the trace donor posed a great challenge. This was due to the inherent uncertainty about the population of alternative donors, the so-called suspect population. We recently proposed to resolve this controversy by systematically favoring the suspect and considering his close male relatives as the suspect population. However, since the mathematical framework developed for this purpose was simulation-based, its practical application turned out increasingly difficult with increasing pedigree size. Here, we present an adaptation of the so-called Elston-Stewart algorithm, originally developed for the linkage analysis of human genetic diseases, to allow calculation of exact match probabilities in a time that scales linearly with pedigree size. The adapted algorithm was implemented in a publicly available software tool, and its correctness was verified by the comparison of its output with the correct, analytical results obtained for selected example pedigrees. The new implementation mostly outperforms the simulation-based solution, albeit with the important exception of Y-STRs present in multiple copies. Given the increasingly prominent role of such multicopy markers in forensic genetics, the complementary use of both approaches appears the most sensible strategy for the time being.