Skip to content

Commit d18bac3

Browse files
authored
Initial commit
0 parents  commit d18bac3

87 files changed

Lines changed: 7305 additions & 0 deletions

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎.github/workflows/ci.yml‎

Lines changed: 75 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,75 @@
1+
name: CI
2+
3+
on:
4+
push:
5+
branches: [main, solution]
6+
pull_request:
7+
workflow_dispatch:
8+
9+
concurrency:
10+
group: ${{ github.workflow }}-${{ github.ref }}
11+
cancel-in-progress: true
12+
13+
jobs:
14+
python:
15+
name: CPU tests (Stages 1-4)
16+
runs-on: ubuntu-latest
17+
timeout-minutes: 30
18+
env:
19+
MUJOCO_GL: osmesa
20+
JAX_PLATFORMS: cpu
21+
steps:
22+
- uses: actions/checkout@v4
23+
24+
- name: Install OSMesa (headless MuJoCo rendering)
25+
run: |
26+
sudo apt-get update
27+
sudo apt-get install -y --no-install-recommends libosmesa6 libgl1 libglfw3
28+
29+
- name: Install uv
30+
uses: astral-sh/setup-uv@v5
31+
with:
32+
enable-cache: true
33+
34+
- name: Sync environment
35+
run: uv sync --locked --dev
36+
37+
- name: Lint
38+
run: uv run ruff check .
39+
40+
- name: Setup check
41+
run: uv run python scripts/check_setup.py
42+
43+
# Student TODOs raise NotImplementedError, which conftest.py reports as
44+
# "not started" (a skip), not a failure. So this job is green on the
45+
# starter `main` for all provided code, and green on `solution` for
46+
# everything.
47+
- name: Fast tests
48+
run: uv run pytest -m "not gpu and not ros and not slow" -q
49+
50+
- name: Slow tests (PPO smoke)
51+
timeout-minutes: 15
52+
run: uv run pytest -m "slow and not gpu and not ros" -q
53+
54+
- name: Progress report
55+
if: always()
56+
run: uv run python scripts/progress.py --no-run
57+
58+
ros2:
59+
name: ROS 2 sim2sim (Stage 5)
60+
runs-on: ubuntu-latest
61+
timeout-minutes: 45
62+
steps:
63+
- uses: actions/checkout@v4
64+
65+
- name: Build the ROS 2 image
66+
run: docker build -t hrc/pup-ros2:jazzy docker/ros2
67+
68+
- name: Run the Stage 5 smoke test
69+
run: |
70+
docker run --rm \
71+
-v "${{ github.workspace }}:/ws" \
72+
-e ROS_DOMAIN_ID=42 \
73+
-w /ws/ros2_ws \
74+
hrc/pup-ros2:jazzy \
75+
/ws/tests/test_05_ros2_smoke.sh

‎.gitignore‎

Lines changed: 23 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,23 @@
1+
.venv/
2+
__pycache__/
3+
*.py[cod]
4+
.pytest_cache/
5+
.ruff_cache/
6+
*.egg-info/
7+
runs/
8+
ros2_ws/build/
9+
ros2_ws/install/
10+
ros2_ws/log/
11+
.DS_Store
12+
.vscode/
13+
.idea/
14+
MUJOCO_LOG.TXT
15+
.pup_progress.json
16+
*.log
17+
18+
# NOTE: results/ is deliberately NOT ignored -- your learning curve, eval JSONs
19+
# and GIFs are part of the submission (see SUBMISSION.md).
20+
21+
# The agent-facing build spec lives on the `solution` branch only; this keeps
22+
# it from being re-committed here when you switch branches.
23+
ONBOARDING_SPEC.md

‎.python-version‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1 @@
1+
3.11

‎README.md‎

Lines changed: 124 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,124 @@
1+
# HRC Software Onboarding — Fall 2026
2+
3+
**Purdue Humanoid Robotics Club · software team**
4+
5+
You are going to build the exact pipeline the club
6+
uses to make NEMO walk: a MuJoCo robot with a PD controller, a vectorized
7+
JAX/MJX environment, a Brax PPO training run, an exported numpy policy, and ROS 2
8+
nodes running that policy against a simulator through the same interfaces the
9+
real hardware uses. The robot is a 12-DoF quadruped called **Pup**. By the end
10+
you will have written a piece of every stage yourself, and you will have a
11+
walking robot to show for it.
12+
13+
![Pup walking under the trained policy](docs/media/pup_walk.gif)
14+
15+
## The pipeline
16+
17+
```mermaid
18+
flowchart LR
19+
S1["<b>Stage 1</b><br/>MuJoCo + PD<br/><i>mujoco, numpy</i>"]
20+
S2["<b>Stage 2</b><br/>JAX for robotics<br/><i>jax</i>"]
21+
S3["<b>Stage 3</b><br/>MJX environment<br/><i>mjx, playground</i>"]
22+
S4["<b>Stage 4</b><br/>Brax PPO + export<br/><i>brax, colab</i>"]
23+
S5["<b>Stage 5</b><br/>ROS 2 sim2sim<br/><i>ros 2, docker</i>"]
24+
S1 --> S2 --> S3 --> S4 --> S5
25+
S1 -. "same PD equation" .-> S5
26+
S3 -. "same 45-dim observation" .-> S5
27+
```
28+
29+
| Stage | Doc | You write | Time |
30+
|---|---|---|---|
31+
| 0 | [Setup](docs/00_setup.md) | — | 0.5–1 h |
32+
| 1 | [MuJoCo and the PD controller](docs/01_mujoco_and_pd.md) | `PDController`, `stand_up`, gain tuning | 2 h |
33+
| 2 | [JAX for robotics](docs/02_jax_for_robotics.md) | three JAX exercises, quaternion maths | 1.5 h |
34+
| 3 | [The MJX environment](docs/03_mjx_environment.md) | observation, termination, commands, `step`, three reward terms | 4 h |
35+
| 4 | [Training with Brax PPO](docs/04_training_with_brax.md) | the training wiring, `export_policy`, `NumpyPolicy` | 2.5 h |
36+
| 5 | [ROS 2 sim2sim](docs/05_ros2_sim2sim.md) | `policy_node`, the launch file | 3 h |
37+
38+
**Total: ~13.5 hours of hands-on work.** This is a generous estimate, and it very possible to finish in much less time.
39+
40+
Also useful: [glossary](docs/glossary.md) ·
41+
[FAQ and troubleshooting](docs/faq_and_troubleshooting.md) ·
42+
[reading packet](docs/reading_packet.md)
43+
44+
## Start here
45+
46+
```bash
47+
# 1. Click "Use this template" on GitHub, then:
48+
git clone https://github.com/<your-username>/onboarding-fall26.git
49+
cd onboarding-fall26
50+
51+
# 2. Install
52+
curl -LsSf https://astral.sh/uv/install.sh | sh # if you don't have uv
53+
uv python install 3.11
54+
uv sync
55+
56+
# 3. Check
57+
uv run python scripts/check_setup.py
58+
uv run python scripts/progress.py
59+
```
60+
61+
Then open [`docs/00_setup.md`](docs/00_setup.md).
62+
63+
## Experience assumptions
64+
65+
We assume you can program in Python, use `git`, and have seen `numpy` arrays
66+
before. We assume **no** background in robot control, reinforcement learning,
67+
JAX, or ROS 2 — every one of those is taught here from zero.
68+
69+
If you have never used the command line for anything beyond `git commit`, this
70+
will be hard but not impossible; budget extra time for Stage 0 and Stage 5, and
71+
ask in Discord early rather than late.
72+
73+
## System requirements
74+
75+
| | Minimum | Notes |
76+
|---|---|---|
77+
| OS | Linux, macOS, or Windows 11 + WSL2 | Linux native is the smoothest |
78+
| Python | 3.11 (3.12 works) | not 3.13 — no `jaxlib` wheel |
79+
| RAM | 8 GB | 16 GB is more comfortable |
80+
| Disk | ~6 GB | ~2 GB of that is the ROS 2 Docker image |
81+
| GPU | not required | Stage 4 uses a free Colab T4; everything else is CPU |
82+
| Docker | required for Stage 5 | or a native ROS 2 Jazzy install |
83+
84+
Stages 1–3 and 5 run on any laptop. Stage 4 is the only one that wants a GPU,
85+
and there is a Colab notebook for it.
86+
87+
88+
## Tests and markers
89+
90+
Your progress bar is:
91+
92+
```bash
93+
uv run python scripts/progress.py # per-stage checklist
94+
uv run python scripts/progress.py --slow # also runs the PPO smoke test
95+
```
96+
97+
Tests report **▷ NOT STARTED** (the function still raises `NotImplementedError`),
98+
**✗ FAILED** (you wrote something and it's wrong), or **✓ PASSED**.
99+
100+
| Marker | Needs | Runs by default? | Command | What it covers |
101+
|---|---|---|---|---|
102+
| *(none)* | CPU only | yes | `pytest` | Stages 1–3 and the Stage 4 export equivalence check |
103+
| `slow` | CPU, under a minute | yes | `pytest -m slow` | the Stage 4 `cpu_smoke` PPO run, end to end |
104+
| `gpu` | an NVIDIA GPU | auto-skipped without one | `pytest -m gpu` | JAX really is on the GPU; vmapped env throughput |
105+
| `ros` | the ROS 2 container | auto-skipped outside it | `pytest -m ros` inside the container | `policy_node`'s observation, ordering, and 50 Hz publish |
106+
107+
Stage 5 as a whole is checked by `tests/test_05_ros2_smoke.sh`, which runs the
108+
`ros` tests and then launches the real stack. It runs inside the container:
109+
110+
```bash
111+
docker compose -f docker/compose.yaml run --rm ros2 /ws/tests/test_05_ros2_smoke.sh
112+
```
113+
114+
The default CPU suite finishes in well under ten minutes.
115+
116+
## Submission
117+
118+
1. Commit your work, plus a `results/` directory containing your learning-curve
119+
PNG, `eval_training.json`, `eval_sim2sim.json`, and your GIFs.
120+
2. Fill in [`SUBMISSION.md`](SUBMISSION.md) — the stage checklist, your pasted
121+
`scripts/progress.py` output, your eval numbers, any escape hatches you used,
122+
the Stage 5 reflection paragraph, and what was hardest.
123+
3. Push to your own `onboarding-fall26` repository.
124+
4. **DM the software lead (Henry Tsay) on Discord once you are done or show it at a meeting.**

‎SUBMISSION.md‎

Lines changed: 91 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,91 @@
1+
# Submission — HRC Software Onboarding Fall 2026
2+
3+
**Name:**
4+
**Discord handle:**
5+
**Repository:**
6+
**Date:**
7+
8+
---
9+
10+
## Stage checklist
11+
12+
- [ ] **Stage 0** — Setup. `scripts/check_setup.py` exits 0.
13+
- [ ] **Stage 1** — MuJoCo + PD controller. Tests green.
14+
- [ ] **Stage 2** — JAX exercises. Tests green.
15+
- [ ] **Stage 3** — MJX environment. Tests green.
16+
- [ ] **Stage 4** — Brax PPO. Trained a policy, exported it, `NumpyPolicy` matches.
17+
- [ ] **Stage 5** — ROS 2 sim2sim. Smoke test green, teleop GIF recorded.
18+
19+
## `scripts/progress.py` output
20+
21+
<details>
22+
<summary>paste the full output here</summary>
23+
24+
```
25+
$ uv run python scripts/progress.py --slow
26+
27+
(paste)
28+
```
29+
30+
</details>
31+
32+
## Stage 4 — training results
33+
34+
**Config used:** (`full` / `t4_fast` / other) &nbsp; **Seed:** &nbsp;
35+
**Where it ran:** (Colab T4 / local GPU / …) &nbsp; **Wall-clock:**
36+
37+
Learning curve: `results/learning_curve.png`
38+
Rollout GIF: `results/training.gif`
39+
40+
```json
41+
(paste results/eval_training.json)
42+
```
43+
44+
Did it meet the acceptance criteria (`"walking_passes": true`)?
45+
46+
## Stage 5 — sim2sim results
47+
48+
Teleop GIF: `results/sim2sim_teleop.gif`
49+
50+
```json
51+
(paste results/eval_sim2sim.json)
52+
```
53+
54+
Measured `/pup/joint_command` rate from `ros2 topic hz`:
55+
56+
## Escape hatches used
57+
58+
- [ ] I used `checkpoints/pup_joystick_flat_reference.npz` for Stage 5 instead
59+
of my own policy.
60+
- [ ] Other (describe):
61+
62+
*(Using one is fine. Not declaring one is not.)*
63+
64+
## Reflection — Stage 5
65+
66+
**List two ways sim2sim can pass while real hardware still fails, and what you
67+
would add to the sim node to catch each.**
68+
69+
>
70+
71+
## What was hardest?
72+
73+
One paragraph. This is will help us improve onboarding.
74+
75+
>
76+
77+
## Time spent
78+
79+
| Stage | Hours |
80+
|---|---|
81+
| 0 Setup | |
82+
| 1 MuJoCo + PD | |
83+
| 2 JAX | |
84+
| 3 MJX env | |
85+
| 4 Brax + export | |
86+
| 5 ROS 2 | |
87+
| **Total** | |
88+
89+
---
90+
91+
**Then DM the software lead (Henry Tsay) on Discord or show during a meeting.**

‎checkpoints/README.md‎

Lines changed: 55 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,55 @@
1+
# Checkpoints — the Stage 5 escape hatch
2+
3+
## What's here
4+
5+
`pup_joystick_flat_reference.npz` — a trained, exported Pup joystick policy, in
6+
the format `pup/train/export.py` produces and `pup/policy/mlp_numpy.py` loads.
7+
8+
## Why it exists
9+
10+
**No stage may be a dead end.** Stage 5 (ROS 2) is worth doing even if Stage 4
11+
(training) did not converge for you — the two teach completely different things,
12+
and getting stuck on one should never cost you the other.
13+
14+
So if your own policy doesn't walk: use this one.
15+
16+
```bash
17+
ros2 launch pup_bringup sim2sim.launch.py \
18+
policy_path:=/ws/checkpoints/pup_joystick_flat_reference.npz
19+
```
20+
21+
That is the default path in the launch file, so this is also what happens if you
22+
pass nothing.
23+
24+
**Say so in `SUBMISSION.md`.** There is a checkbox for it. Declaring an escape
25+
hatch costs you nothing; quietly pretending you trained it costs you the
26+
reviewer's trust, and they can tell — your `results/` has to contain a learning
27+
curve and an eval JSON from your own run either way.
28+
29+
## Using your own instead
30+
31+
```bash
32+
uv run python -m pup.train.export \
33+
--checkpoint runs/colab/policy.pkl \
34+
--out results/my_policy.npz
35+
```
36+
37+
Then point `policy_path` at it. Commit it — it is a couple of hundred kilobytes.
38+
39+
## What's inside the file
40+
41+
```python
42+
import numpy as np
43+
archive = np.load("checkpoints/pup_joystick_flat_reference.npz")
44+
print(archive.files)
45+
print(str(archive["obs_layout"]))
46+
```
47+
48+
`obs_mean` / `obs_std` (the observation normalizer), `kernel_i` / `bias_i` for
49+
each MLP layer in forward order, and metadata: `n_layers`, `obs_size`,
50+
`action_size`, `action_scale`, `default_pose`, `hidden_activation`,
51+
`obs_layout`.
52+
53+
`obs_layout` is the important one. It is the authoritative statement of what the
54+
45 numbers mean, travelling inside the file with the weights, so that a policy
55+
and a robot can never quietly disagree about it.
751 KB
Binary file not shown.

0 commit comments

Comments
 (0)