Taktile Labs

Advancing the frontier of AI in financial decisions.

Taktile Labs is an applied research group focused on advancing the frontier of AI in financial decisions: exploring, testing and proving new techniques across the areas that matter in banking and insurance, and sharing the evidence.

Bridging the gap between frontier AI and regulated deployment.

Financial services will spend $97B on AI by 2027. Yet the distance between what general-purpose AI can do and what regulated institutions can reliably deploy remains vast. Taktile Labs exists to close this gap.

Focus on AI adoption within critical decision-making.

Every day, financial institutions run on a current of critical decisions, from determining if a transaction is fraudulent to calculating the amount of risk to accept when underwriting a loan. At Taktile Labs, we research what it takes to deploy AI in this high-stakes decision-making where errors aren’t an option.

Powered by high-quality, realistic data.

We work closely with development partners and industry experts to create high-quality evaluation data sets, which we use to assess AI performance in financial services-specific contexts.


What we’re working on.

We’re pursuing five research tracks focused on the most important requirements for trustworthy AI in regulated financial institutions.

01

Evaluations & Benchmarking

In collaboration with our development partners, we use real-world data to build trusted benchmarks for model performance in core financial services use cases. Every benchmark is designed around the KPIs business teams care about most: accuracy, cost per decision, and latency.

View UBO-Bench →
02

Human-Agent Design Patterns

Not every decision should be fully automated, and not every decision requires a human to intervene. We pursue balanced research that helps teams unlock the benefits of AI in complex decision-making while preserving the value of human judgment and engagement.

03

Governance, Risk & Compliance

Agentic systems based on LLMs are stochastic and hard to inspect, challenging assumptions behind SR 11-7 and traditional model risk management. We work with institutions and regulators to clarify what responsible adoption looks like in practice.

04

Foundation Models for Financial Data

Foundation models are trained on public text, but financial decisioning runs on data with rich sequential and relational structure. We explore transformer-based architectures purpose-built for common data structures in financial services.

05

Hybrid Decision Architectures

The most effective AI systems will be built using hybrid architectures. We help financial institutions navigate AI hype with clarity and choose the right tool for each task while balancing cost, risk, and performance.

Discover our benchmark for adverse media research.

KYBench evaluates how well agentic AI systems can investigate a business, weigh the evidence they find, and judge regulatory red flags in realistic Know Your Business reviews. Built with real businesses annotated by expert compliance practitioners.

token spendfalse positive reviewmissed findings, 30 min each
GPT-6 Astra
$3.46
Claude Opus 5
$3.80
GPT-5.6 Sol
$3.81
Claude Sonnet 5
$3.93
GPT-5.6 Luna
$3.94
GPT-5.6 Terra
$4.36
Claude Opus 4.6
$4.54
Gemini 3.1 Pro
$4.94
Claude Opus 4.7
$5.11
GPT-5.4
$5.20

Total economic cost per business screened

September 2026 update. All 25 models and methodology on the benchmark page.

Meet our team.

Taktile Labs is powered by consistent collaboration between a dedicated internal research team and an external Research Council and Advisory Board.

Advisory Board

Parag AgrawalTom GlocerHarald SchneiderKarim LakhaniRobin GreenwoodTina ReichJill Zucker Sheckman

Research Council

Fagner AbreuYoni CohenBen LiebaldDaniel MeyerJonas NelleMikey ShulmanPieter ViljoenMichael Zambrano

Research Team

David AhnMaximilian Eber, PhDNico KleesAlexia PastréFabian Peters, PhDRobin Raymond, PhD

Browse our published work.

Benchmarks, technical reports, and research.

BenchmarkSeptember 2026

UBO-Bench: Evaluating AI Agents for Beneficial Ownership Checks

Moritz GeistDavid AhnMaximilian Eber, PhD

Moritz Geist, David Ahn, Maximilian Eber, PhD

A benchmark of AI agents that trace ownership structures for beneficial ownership checks on UK and German registers. Measures tracing accuracy across frontier models, importance of country specific configuration, and how much of the work agents complete on their own.

BenchmarkJune 2026

PIBench: Prompt Injection Resistance in Agentic Underwriting

Koen RoelofsJakob SchmittMaximilian Eber, PhD

Koen Roelofs, Jakob Schmitt, Maximilian Eber, PhD

The first benchmark of prompt-injection resistance for agentic underwriting. Measures defense success across 16 frontier models, three providers, and five attack vectors — with and without untrusted-content tagging.

BenchmarkApril 2026 · updated Sep 2026

KYBench: Evaluating AI Agents for Adverse Media Research

David AhnMaximilian Eber, PhDSahith Jagarlamudi

David Ahn, Maximilian Eber, PhD, Sahith Jagarlamudi

The first public benchmark of AI-driven adverse media investigation. Evaluates detection accuracy, evidence quality, reliability across agent runs, and cost efficiency across frontier models.

BenchmarkMarch 2026

FinSpread-Bench: Evaluating Agentic AI for Financial Spreading

Nico KleesMaximilian Eber, PhD

Nico Klees, Maximilian Eber, PhD

The first public benchmark for agentic financial document processing. Evaluates extraction accuracy, cross-document reasoning, calculation correctness, and structured output quality across seven frontier models. Built on anonymized production data from financial institutions.

PaperMarch 2026

AI in AML: A guide to governance and implementation

Dustin EatonMaximilian Eber, PhD

Dustin Eaton, Maximilian Eber, PhD

Why AML teams must now apply model risk management standards to AI systems. Published in ACAMS Today, exploring how regulators are extending MRM frameworks to AI deployed in compliance functions — and what institutions need to do to prepare.