← All projects
03Aug 2026

Agentic Fraud Investigation System

FastAPIPythonXGBoostReactAnthropic APIPostgreSQL

Overview

A fraud detection microservice architecture combining high-throughput ML scoring with an autonomous multi-tool agent. High-risk transaction events trigger an LLM agent loop that executes geo/IP verification, account graph queries, and counterparty reputation tools to synthesize audit-ready investigation dossiers for human review.

Approach

  • Trained an XGBoost & IsolationForest anomaly detection pipeline on transaction data enriched with device IDs, IP subnets, and counterparty graphs.
  • Built 4 RESTful investigation tool integrations allowing the autonomous agent to query geolocation, historical velocity, shared device networks, and account history.
  • Constructed a human-in-the-loop audit interface ensuring every automated recommendation requires human authorization, backed by an LLM-as-a-judge evaluation harness.
XGBoost + IsolationForest → risk score
Geo/IP
Customer history
Device network
Reputation
Human reviewer: approve / reject / re-investigate
Transaction → risk model → agent investigation → human review
Fraud Investigation Console live dashboard
Fraud Investigation Console live dashboard

Results

Metrics — sourced from the repo (README / model card / test output)

Classifier test PR-AUC / ROC-AUC
1.0000 / 1.0000

PaySim balance-consistency leakage, documented in the model card — not a production-transferable number

Precision / recall @ threshold 0.525
1.000 / 0.970
Companion anomaly signal (IsolationForest ROC-AUC)
0.840
Investigation tool integrations
4 (geo/IP, history, device graph, reputation)
Gold-standard eval set
20 cases, LLM-judge scored
Sync investigation latency
20–40s

flagged in-repo as needing a background queue for production

Eliminated black-box decision output by backing every risk flag with reproducible API tool trace logs and structured evidence reports. Fully deployed full-stack web application; the base classifier's near-perfect test metrics are a documented artifact of the benchmark dataset, not a production claim — see the metrics below and the repo's model card.

Impact