Agentic Fraud Investigation System
Overview
A fraud detection microservice architecture combining high-throughput ML scoring with an autonomous multi-tool agent. High-risk transaction events trigger an LLM agent loop that executes geo/IP verification, account graph queries, and counterparty reputation tools to synthesize audit-ready investigation dossiers for human review.
Approach
- Trained an XGBoost & IsolationForest anomaly detection pipeline on transaction data enriched with device IDs, IP subnets, and counterparty graphs.
- Built 4 RESTful investigation tool integrations allowing the autonomous agent to query geolocation, historical velocity, shared device networks, and account history.
- Constructed a human-in-the-loop audit interface ensuring every automated recommendation requires human authorization, backed by an LLM-as-a-judge evaluation harness.

Results
Metrics — sourced from the repo (README / model card / test output)
- Classifier test PR-AUC / ROC-AUC
- 1.0000 / 1.0000
- Precision / recall @ threshold 0.525
- 1.000 / 0.970
- Companion anomaly signal (IsolationForest ROC-AUC)
- 0.840
- Investigation tool integrations
- 4 (geo/IP, history, device graph, reputation)
- Gold-standard eval set
- 20 cases, LLM-judge scored
- Sync investigation latency
- 20–40s
PaySim balance-consistency leakage, documented in the model card — not a production-transferable number
flagged in-repo as needing a background queue for production
Eliminated black-box decision output by backing every risk flag with reproducible API tool trace logs and structured evidence reports. Fully deployed full-stack web application; the base classifier's near-perfect test metrics are a documented artifact of the benchmark dataset, not a production claim — see the metrics below and the repo's model card.