Portrait of Elham Daneshmand

Elham Daneshmand

PhD candidate, McGill University & Mila · Montreal, Canada

I am a final-year PhD candidate in Computer Science at McGill University and Mila (GPA 4.00), advised by Glen Berseth and Hsiu-Chin Lin.

My research is on post-training: teaching a pretrained model a new skill or a new setting without losing what it already knows. I study it on robot policies, where forgetting and unsafe exploration have physical costs. I began with model predictive control for bipeds, then safe reinforcement learning and on-robot fine-tuning for quadrupeds. More recently, I have turned to post-training methods from large language models, such as low-rank adaptation, KL-regularized fine-tuning and GRPO, and my current work is on policies up to a 7B vision-language-action model. The same questions drive LLM post-training: what to update, how far to move from the base model, and how to add a capability without losing the others.

My work spans humanoid whole-body tracking and loco-manipulation on the Unitree G1, and safe fine-tuning on a physical Unitree Go2 as a visiting researcher at the Technical University of Munich, supported by Majid Khadiv. Before my PhD, I completed a B.Sc. in Computer Science at Amirkabir University of Technology and worked on model predictive control for torque-controlled bipeds at the Max Planck Institute for Intelligent Systems, in the group of Ludovic Righetti (NYU).

McGill undergraduates interested in a robot-learning project are welcome to email me with a short note on their interests.

I am open to research internships in 2027.

  • post-training
  • sim-to-real transfer
  • RL post-training
  • humanoid whole-body control
  • safe reinforcement learning
  • parameter-efficient fine-tuning (LoRA)
Fine-tuning on a real Unitree Go2 Push recovery on the biped Bolt Uneven ground on the biped Bolt

News

Publications

Simulated Unitree G1 humanoid tracking a reference motion
Simulated Unitree G1 motion tracking with the pretrained policy, one of the study's settings.

Fine-Tuning Methods for Post-Training Robot Policies

E. Daneshmand et al.

In preparation

When a pretrained policy has to learn something new, which fine-tuning method keeps what it already knows? A head-to-head comparison, from humanoid motion tracking to a 7B vision-language-action model.

Illustration: the safe region grows during training on a curriculum; outside it, a recovery controller acts.

SafeExplorer: An Unbiased Mixed-Policy Gradient for Reinforcement Learning with Recovery Interventions

E. Daneshmand, M. Khadiv, G. Berseth, H.-C. Lin

Transactions on Machine Learning Research (TMLR), 2026 (accepted)

Keeping a robot safe while it learns means letting a recovery controller take over, and that quietly biases every policy-gradient update. SafeExplorer removes the bias while keeping the safety net.

Physical Unitree Go2 jump: the pre-trained policy, on-robot fine-tuning (sped up), then the fine-tuned policy.

SLowRL: Safe Low-Rank Adaptation Reinforcement Learning for Locomotion

E. Daneshmand, S. Omar, G. Berseth, M. Khadiv, H.-C. Lin

Under review, IEEE-RAS Humanoids 2026 · Workshop poster, RLC 2026

Can a policy trained in simulation be fine-tuned on the real robot, safely, by adapting only a low-rank slice of its weights? SLowRL brings LoRA, the adapter behind efficient LLM fine-tuning, to on-robot reinforcement learning.

The Bolt biped walking on uneven ground
The real biped Bolt on uneven ground.

Variable Horizon MPC with Swing Foot Dynamics for Bipedal Walking Control

E. Daneshmand, M. Khadiv, F. Grimminger, L. Righetti

IEEE Robotics and Automation Letters (RA-L), 2021 · presented at ICRA 2021

Robot video featured in IEEE Spectrum Video Friday, October 2020

Where and when should a biped put its next foot after a push? A two-level MPC decides both, and keeps the real biped Bolt on its feet.

Simulation snapshots of the Bolt biped walking and running
Walking and running gaits for Bolt (simulation).

A Unified Framework for Walking and Running of Bipedal Robots

M. Ghoddousi Boroujeni*, E. Daneshmand*, L. Righetti, M. Khadiv (*equal contribution)

International Conference on Advanced Robotics (ICAR), 2021

Walking and running usually need different models. Relaxing one assumption of the classic pendulum model, a fixed centre-of-mass height, lets one framework generate both.

Koala Science logo

Koala Science: A Platform for AI Reviewers in the Wild

The Koala Science team and five top-ranked competition participants, including E. Daneshmand

NeurIPS 2026 Workshop on AI-Native Academia (accepted)

Can AI agents review papers? This platform ran reviewing agents on real ICML 2026 submissions and scored their verdicts against the actual decisions. I co-authored it as one of the five top-ranked competition participants who helped test the platform.

Other projects

Experience