VLA / Embodied AI Algorithm Intern Candidate · Shenzhen, China
M.S. student in Electronic Information at the University of Chinese Academy of Sciences / Shenzhen Institutes of Advanced Technology (SIAT) · 2025–2028
I work at the intersection of robot data, vision-language-action policies, simulation evaluation, and reliable deployment. My preferred workflow is measurable and reproducible: define the data contract, validate the rollout, inspect failure cases, then optimize the system bottleneck.
- Built the data path Raw → LeRobot v2.1 → DexData → training view and adapted a mock-compatible worker interface for 30 tasks.
- Processed 32,939 episodes / 45,827,921 frames and isolated 15 anomalous episodes.
- Added SHA-256 manifests so that copied artifacts can be checked before training or evaluation.
- Supported ×ばつA100 80 GB BF16 distributed training with ZeRO-3 and gradient checkpointing; retained resume validation records.
- Adapted deployment interfaces, including camera reordering, 14D↔32D action mapping, joint-delta / gripper-absolute decoding, and fail-closed safety gates.
- Verified 35/35 CPU checks, an A100 BF16 forward path, and the official mock worker.
- Evaluated an InternVL3-1B + TinyVLA configuration on a local WBCD-2026 branch/configuration.
- Ran 10 tasks ×ばつ 100 local rollouts with selected rollout videos retained for inspection.
- Local results:
open_microwave87/100,move_can_pot3/100, and 0/100 on the other eight tasks. - Reduced the local evaluation workflow from roughly 8 hours to 45 minutes through batching and execution tooling.
The RoboTwin numbers above are local simulation evidence. I do not present them as an official leaderboard score(官方榜单成绩). The RoboChallenge record currently includes an official mock and deployment checks, but not a real W1 success rate or rank(真实 W1 成功率或排名).
- robotwin-evaluation-tools — local rollout summaries, integrity manifests, and evaluation scaffolding.
- robot-data-pipeline — schema-first dataset conversion and quality-control framework.
- vla-paper-reading-notes — structured notes on RT-1, RT-2, OpenVLA, ACT, Diffusion Policy, π0, and RTC.
The repositories contain lightweight, reproducible scaffolds. Private datasets, credentials, internal URLs, model weights, and real-robot logs are intentionally excluded.
Vision-Language-Action · robot imitation learning · action chunking · diffusion / flow-matching policies · LeRobot · PyTorch · distributed training · data quality · simulation evaluation · ROS 2 / C++ · deployment safety
- Make the contract explicit: observation timestamps, camera order, action dimensions, units, normalization, padding, and masks.
- Make failures countable: isolate bad episodes, preserve manifests, and report per-task metrics rather than a single aggregate number.
- Measure the closed loop: include preprocessing, model, postprocessing, transport, controller, P50/P95 latency, and safety fallbacks.
- State the evidence boundary: distinguish a paper I read, a component I implemented, a local simulation result, and a real-robot result.
- Reproducing small, public ACT and Diffusion Policy baselines in simulation.
- Strengthening ROS 2, C++、FK/IK、Jacobian、PID/MPC and sim-to-real debugging.
- Building public experiment reports with configs, checksums, failure taxonomies, and latency measurements.
- GitHub: Airgo0911
- Location: Shenzhen, China
- For collaboration or internship discussions, please open an issue or use the contact channel listed on my current resume.
This profile is a concise portfolio summary. Please verify current availability and project status in the linked repositories and resume.