CoFiT, a method for fine-tuning humanoid tracking policies around runtime safety filters, reduced violation time across diverse constraint scenes and on Unitree G1 hardware, first reported by Arxiv. The approach accounts for filtering’s changes to executed actions and policy-induced state distributions, with tests showing lower violation time and smaller safety-filter corrections.
Safe whole-body motion is essential to deploy humanoid robots in unstructured environments. Humanoid control commonly separates reference specification from execution: a planner, teleoperator or motion generator supplies a reference, while a reinforcement-learning policy tracks it through dynamically feasible whole-body control.
Runtime safety filters can intervene on tracker outputs to enforce newly introduced constraints. Those interventions alter executed actions and the state distribution induced by the policy; treating the tracking policy and filter independently produces dynamics, objective and information mismatches.
Case studies isolate those three mismatch types and examine their root causes. CoFiT, short for Constrained Filter-aware Tuning, fine-tunes pretrained trackers with the policy-filter interaction in view.
Across diverse constraint scenes, CoFiT reduced violation time relative to filter-only training by 91% on TWIST2 and 21% on SONIC, while requiring smaller safety-filter corrections. The reported results cover violation time, filter-correction size and operator intervention.
On Unitree G1 hardware, CoFiT reduced violation time by 83% for TWIST2 and completed every trial without operator intervention. An operator stop was required in 50% of baseline trials.