TQBTHEQUANTBATEMAN
TQB/ learn/ frontier/ deep hedgingEN · DARK
Frontierresearchresearch

Deep Hedging

Learn hedging policies under frictions and non-quadratic objectives.

Reviewed 2026-08-10TheQuantBateman ResearchReading note
01Intuition

State the empirical or computational motivation.

Optimise the trading policy directly when transaction costs and constraints break textbook replication.

ONE-LINE DEFINITION

Learn hedging policies under frictions and non-quadratic objectives.

02Mathematics

Expose the proposed mathematical object.

min⁡π  ρ(hedging P&Lπ)\min_\pi \; \rho(\text{hedging P\&L}_\pi)
Notation and units

Decimal rates and volatilities, year-fraction time and continuous compounding unless stated otherwise.

03Assumptions

Separate evidence from modelling choice.

01

The proposed method is compared with an established baseline on held-out scenarios.

02

Parameter uncertainty and extrapolation are reported rather than hidden by one fit metric.

03

Production use requires independent validation, monitoring and a documented fallback.

“An unstated convention is a future reconciliation break.”— THEQUANTBATEMAN
04Market use

Define a falsifiable validation target.

Active research and selective experimentation, not a universal replacement for desk risk systems.

Intuition→Mathematics→Implementation→Desk risk
05Desk view
FRONT OFFICE VIEW

Treat governance as part of the method.

Keep an established Frontier baseline beside the new method and define the scenario in which the fallback takes control.

Ask Bateman about this topic →
06Related

Trace the nearest established baseline.