← All projects
04Aug 2026

Oil Market Manipulation Analysis

PythonXGBoostARIMAReinforcement LearningFastAPI

Overview

A modular analytical software testbed running 5 concurrent evaluation engines — SARIMAX time-series forecasting, XGBoost regression, Granger causality testing, PPO reinforcement learning, and unsupervised anomaly detection — over multi-asset market and subsidy data streams.

Approach

  • Designed a modular data engine processing synthetic market indicators timed to historical macroeconomic shock points (OPEC cuts, sanctions, demand shifts).
  • Engineered Granger causality testing to differentiate statistical lead-lag relationship directions from simple correlation spikes.
  • Built a comparative execution framework evaluating supervised classification, unsupervised clustering, and PPO reinforcement learning agents against identical stream inputs.
oil subsidy →manipulation index ↑
Manipulation index vs. oil subsidy — flagged outliers in gold
1
2
3
4
5
6
7
8
9
10
11
12
lag 1 — p<0.0001lag 12 — p=0.117

Significant through lag 8 (p<0.01). Intervention leads hidden subsidy by up to eight months.

Granger causality: govt intervention → hidden subsidy, by lag (months)
manipulation index →hidden subsidy ↑
Hidden subsidy (symlog) vs. manipulation index
2022 sanctions— govt intervention— manipulation index
Govt intervention vs. manipulation index, 2015–2023 (38-country mean)

Results

Metrics — sourced from the repo (README / model card / test output)

Evaluation engines
5 (SARIMAX, regression sweep, Granger, anomaly detection, PPO)
Regression model families swept
7 (Linear, Ridge, KNN, Decision Tree, Random Forest, XGBoost, LightGBM)
Real macro events anchoring synthetic data
4 (2016 OPEC cuts, 2018 sanctions, 2020 COVID demand collapse, 2022 Russia sanctions)
Deployment status
not yet deployed

next step per repo: validate against a labeled or real-world dataset

Achieved multi-model alignment across independent algorithmic lenses while maintaining strict validation boundaries before real-world dataset deployment. The dataset is explicitly synthetic — documented as such in the repo rather than presented as observed market data — so the metrics below describe engineering scope, not a deployed-accuracy claim.

Impact