| title | Fireworks Serverless RL |
|---|---|
| description | Train a Fireworks LoRA adapter on rollouts from a HUD environment, with grouped rewards, checkpoints, evaluation, and a path to hosted agent tasks. |
| icon | /logo/fireworks.svg |
A model trainer needs more than prompts. It needs a repeatable interaction, an execution environment, and a reward that measures whether the model completed the task. HUD packages those pieces together, so the same environment can be used to evaluate a model, collect training rollouts, and verify whether training improved it.
This cookbook connects that environment layer to Fireworks Serverless Training. The environment author defines the task and reward. HUD executes the environment and grader. Fireworks samples and trains the model. The cookbook handles orchestration, advantages, and datum construction between them.
The complete project includes the environment, training loop, tests, and dependency configuration.%%{init: {'flowchart': {'padding': 24, 'nodeSpacing': 48, 'rankSpacing': 56}, 'themeVariables': {'fontSize': '17px'}}}%%
flowchart TD
subgraph hud["HUD"]
direction TB
taskset["Taskset"]
env["Environment<br/>task · tools · grader"]
taskset --> env
end
subgraph cookbook["Cookbook"]
direction TB
agent["FireworksAgent"]
batch["Advantages → datums"]
end
subgraph fireworks["Fireworks"]
direction TB
sampler["Sampler<br/>LoRA snapshot"]
trainer["Serverless Training"]
end
env <-->|prompt · response| agent
agent <-->|sample request · response| sampler
env -->|graded Runs| batch
batch -->|training datums| trainer
trainer -->|save updated snapshot| sampler
classDef hudNode fill:#eaf4ff,stroke:#2563eb,stroke-width:2px,color:#172033;
classDef cookbookNode fill:#f3efff,stroke:#7c3aed,stroke-width:2px,color:#24143d;
classDef fireworksNode fill:#fff0e5,stroke:#f97316,stroke-width:2px,color:#3d1f0c;
class taskset,env hudNode;
class agent,batch cookbookNode;
class sampler,trainer fireworksNode;
style hud fill:#f8fbff,stroke:#60a5fa,stroke-width:2px,color:#172033;
style cookbook fill:#fbf9ff,stroke:#a78bfa,stroke-width:2px,color:#24143d;
style fireworks fill:#fffaf5,stroke:#fb923c,stroke-width:2px,color:#3d1f0c;
The boundary between the two systems is explicit:
| HUD | Fireworks |
|---|---|
| Defines tasks, tools, state, and graders | Hosts the base model and LoRA adapter |
| Executes rollouts locally or on hosted runtimes | Samples from an adapter snapshot |
| Returns rewards and traces | Runs forward, backward, and optimizer operations |
| Groups repeated attempts at the same task | Stores training state and sampler checkpoints |
The cookbook supplies the FireworksAgent adapter, computes group-relative advantages, and converts
completed Runs into Fireworks training datums.
The example uses multiplication because it is small and easy to verify. HUD also represents environments with browsers, code execution, games, APIs, and stateful tools; the adapter changes needed for those environments are covered below.
Install Git, Python 3.11 or 3.12, and
uv. Then clone the SDK and enter the
cookbook directory:
git clone https://github.com/hud-evals/hud-python.git
cd hud-python/cookbooks/fireworks-rl-training
uv syncRun the remaining commands from hud-python/cookbooks/fireworks-rl-training.
Set a Fireworks API key with Serverless Training access:
export FIREWORKS_API_KEY="fw_..."The default configuration uses Qwen 3.8 27B (accounts/fireworks/models/qwen3p8-27b)
with tokenizer Qwen/Qwen3.8-27B. Fireworks deprecated the Qwen 3.5 9B and Qwen 3.6 27B
Serverless Training pools; see the
Fireworks changelog.
If you change the model, also provide the matching tokenizer and renderer:
uv run train.py \
--base-model "<fireworks-model>" \
--tokenizer-model "<hugging-face-tokenizer>" \
--renderer "<renderer-name>" \
--calibratePrompt rendering happens on the client, so the base model, tokenizer, and renderer must agree. The
default renderer, qwen3_8_disable_thinking from tinker-cookbook>=0.5.7, disables thinking mode.
This leaves enough of the generation budget for the model to emit the final answer that the
grader reads.
The environment is a normal HUD environment. The task yields a prompt, receives the model response,
and returns an EvaluationResult:
import re
from hud import Environment
from hud.graders import EvaluationResult
env = Environment(name="fireworks-arithmetic")
def grade_final_integer(answer: object, expected: int) -> EvaluationResult:
text = (answer if isinstance(answer, str) else str(answer)).strip()
final = text.splitlines()[-1].strip() if text else ""
final = (
re.sub(r"\\boxed\{\s*([^{}]+?)\s*\}", r"\1", final.rstrip(".")).strip("$*` \t").rstrip(".")
)
got = (
int(final.replace(",", ""))
if re.fullmatch(r"[+-]?(?:\d+|\d{1,3}(?:,\d{3})+)", final)
else None
)
return EvaluationResult(
reward=1.0 if got == expected else 0.0,
content=text,
info={"expected": expected, "got": got},
)
@env.template()
async def multiply(a: int, b: int):
answer = yield (
f"What is {a} * {b}? Work it out, then put the final integer "
"on its own line at the end of your answer."
)
yield grade_final_integer(answer, a * b)The grader belongs to the environment rather than the trainer. That keeps the task definition
portable: the same multiply task can be run in an evaluation, used with another model, or sent to
a different training backend without rewriting the reward.
Group-relative training compares repeated attempts at the same task. If every rollout in a group receives the same reward, the normalized advantages are zero and the group cannot update the model.
Calibration mode samples from the initial adapter and reports the reward spread without taking an optimizer step:
uv run train.py \
--calibrate \
--tasks-per-step 6 \
--group-size 4 \
--max-tokens 2048 \
--debug-samples 4| Metric | Interpretation |
|---|---|
reward_mean |
Overall task difficulty for the current model |
within_group_reward_std |
Training signal available within repeated attempts |
Positive within_group_reward_std confirms reward variation in at least one group. It does not
prove that the grader is correct. --debug-samples prints responses with their rewards and token
counts; verify that better answers receive higher rewards. If all groups are correct, increase the
operand range with --min-a, --max-a, --min-b, and --max-b. If all groups are incorrect,
reduce the range or increase --max-tokens.
The default uses four-digit operands and a 2,048-token budget. Three-digit tasks were nearly
saturated on Qwen 3.8 27B, while smaller budgets cut off many answers. Inspect the full
--debug-samples responses and token counts so truncation is not mistaken for arithmetic difficulty.
After calibration shows useful reward spread, run one bounded training step:
uv run train.py \
--steps 1 \
--tasks-per-step 2 \
--group-size 4 \
--max-tokens 2048 \
--eval-tasks 4 \
--require-updateGroups with identical rewards cannot update the model. --require-update makes the command fail
instead of silently skipping the optimizer. A successful run verifies authentication, sampling,
HUD grading, one policy-gradient update, checkpoint creation, and held-out evaluation.
Sampling or grading errors fail the command in every phase, including calibration and evaluation.
The default command requests 30 steps × 8 task groups × 8 attempts, or 1,920 training rollouts, followed by 16 evaluation rollouts:
uv run train.pyEach rollout can generate up to 2,048 tokens and incurs Fireworks usage. Charges include prompt prefill, sampled output, and training tokens at the selected model's serverless rates.
Each step saves the current adapter for sampling, runs the HUD taskset, converts graded runs into training datums, and applies an update when at least one group has reward variation:
Conceptual excerpt from train.py; this is not standalone code.
snapshot = training_client.save_weights_for_sampler(
f"policy-{step:04d}"
).result()
sampler = service.create_sampling_client(
model_path=snapshot.path,
tokenizer=tokenizer,
)
job = await taskset.run(
agent,
runtime=runtime,
group=group_size,
)
datums, kept_groups = make_training_batch(job.runs)
if datums:
training_client.forward_backward(
datums,
"importance_sampling",
).result()
training_client.optim_step(adam).result()FireworksAgent adapts the sampler to HUD's Agent interface. It records the prompt tokens,
generated tokens, and sampling logprobs on each Run. make_training_batch then:
- Groups runs that attempted the same task.
- Standardizes rewards within each group.
- Drops groups with no reward variation.
- Assigns the group advantage to each generated token.
- Builds the datums expected by the Fireworks training client.
Metrics are written to runs/fireworks-serverless/metrics.jsonl. Each row includes mean reward,
within-group reward spread, valid rollout count, retained groups, training datums, whether an update
was applied, loss, snapshot, and step duration.
Sampling and training checkpoints serve different purposes:
| Checkpoint | Contains | Use |
|---|---|---|
policy-*, final |
Adapter weights | Bind an in-session sampler or promote the adapter |
state-*, final-state |
Adapter weights and optimizer state | Resume training |
Resume from a fully qualified training checkpoint:
uv run train.py --resume-from "<account>/<run-id>/state-0005"Resume restores the checkpoint's original base model. Start a new run to migrate from an older Qwen model to Qwen 3.8; resuming does not convert the adapter. A supported non-default checkpoint requires its matching tokenizer and renderer.
The resumed session creates a new Fireworks run and retains the optimizer state. The script prints
the final Sampler checkpoint path. Sampler checkpoints are session-scoped, so use that path to
identify and
promote the checkpoint
before the session is removed.
For a local one-turn environment, pass its tasks and environment files:
uv run train.py \
--tasks-file "../my-environment/tasks.py" \
--env-path "../my-environment/env.py" \
--calibrate \
--tasks-per-step 6 \
--group-size 6Calibration needs at least --tasks-per-step tasks. A training run needs
--tasks-per-step + --eval-tasks; the script shuffles them deterministically and creates disjoint
training and evaluation subsets.
For a hosted taskset, first
deploy the environment, sync the
tasks, and set HUD_API_KEY. Then pass the taskset name or id:
uv run train.py \
--taskset "my-taskset" \
--calibrate \
--tasks-per-step 6 \
--group-size 6A hosted run executes the environment and grader on HUD while the cookbook's FireworksAgent
continues to call the Fireworks sampler from the training process.
The included FireworksAgent and batch builder support one generated assistant response per Run.
Use a different agent adapter and batch builder for tool-using or multi-turn environments so every
trainable assistant turn is recorded and user and tool-result tokens remain masked.