← Kaidi Zha

AIDER: Taming Aleatoric Impulse in Off-Policy Reinforcement Learning

Reinforcement Learning & Control, course project (team of 8)  ·  Spring 2026

Overview

Off-policy reinforcement learning can suffer from a transient surge in value overestimation during early training. AIDER addresses this with Direct Evidential Regression (DEAR), which uses an EMA anchor to learn aleatoric value uncertainty directly, then combines it with epistemic uncertainty for pessimistic target evaluation and dynamic exploration.

Method

DEAR treats an estimator as a Gaussian random variable and regresses its variance against a stable EMA target, avoiding heuristic annealing. AIDER embeds this evidential head into the critic so that the learned uncertainty supplies both a lower-confidence-bound target penalty and an upper-confidence-bound exploration bonus.

Results

On dynamic DeepMind Control Suite locomotion tasks, AIDER improves over its closest predecessor, with the largest gains on unstable high-dimensional tasks such as Humanoid-run and Dog-run. On a four-wheel-drive vehicle tracking task, the learned policy follows the sinusoidal reference speed closely while producing coordinated torque commands.

Training curves of model-free algorithms on MuJoCo and DMC benchmarks
Benchmark training curves — AIDER leads on dynamic locomotion tasks such as Humanoid-run and Dog-run.
Vehicle speed tracking on a representative test episode
Vehicle deployment — actual speed tracks the sinusoidal reference with coordinated four-wheel torque commands.
DEAR vs. DER final fit comparison on the toy regression task
Toy regression — DEAR tracks the true curve with a tighter predictive band than DER.
Uncertainty decomposition of DEAR and DER
Uncertainty decomposition — DEAR remains compact in-domain and rises smoothly out-of-distribution.

← Back to home