Humanoid and Bipedal Locomotion
PPO locomotion for the Hector biped and motion-retargeting and behavior-cloning pipelines for Unitree G1 whole-body imitation.
Research
I study how robots can continue to move and act safely when actuators, terrain, and observations depart from nominal conditions. My work spans reinforcement learning, safety-critical control, real-world robot integration, and efficient multimodal inference.
Actuator weakness, lock-up, and loss of torque change the forces and motions a robot can physically realize. I develop fault-conditioned locomotion policies and safety filters that account for these changes instead of relying on explicit fault detection alone.
The current framework combines curriculum-based reinforcement learning with actuator torque-speed envelopes, floating-base dynamics, and control barrier functions. A quadratic program minimally modifies commands while enforcing feasible contact forces, joint limits, and roll and pitch safety conditions on the Unitree Go2.
I worked on training-free visual-token pruning for autoregressive and diffusion vision-language models. HIVTP uses middle-layer attention and hierarchical global-local selection to retain fine-grained information. RedVTP uses attention from still-masked response tokens to prune visual tokens after the first diffusion inference step.
Across the evaluated diffusion VLMs, RedVTP reduced inference latency by up to 64.97% and improved token-generation throughput by up to 186% while preserving—and sometimes improving—benchmark accuracy.
Projects that shaped my experience in robot learning, systems optimization, and sequential decision-making.
PPO locomotion for the Hector biped and motion-retargeting and behavior-cloning pipelines for Unitree G1 whole-body imitation.
Behavior cloning and Soft Actor-Critic for cold-start-aware scheduling under dynamic edge-resource constraints.
PPO, graph attention, and pointer networks for generating strong initial columns in aircraft-recovery optimization.