Skip to content

docs: freeze M for certified-reward RL Phase-1/pilot - #334

Merged
abrichr merged 1 commit into
mainfrom
agent/certified-reward-m-freeze
Sep 2, 2026
Merged

docs: freeze M for certified-reward RL Phase-1/pilot#334
abrichr merged 1 commit into
mainfrom
agent/certified-reward-m-freeze

Conversation

@abrichr

@abrichr abrichr commented Sep 2, 2026

Copy link
Copy Markdown
Member

The certified-reward RL pre-reg (section 3.5) says freeze M before the first training step. This PR is that freeze for the Phase-1/pilot.

It hash-pins:

  • checkpoint Qwen/Qwen2.5-VL-3B-Instruct at revision 66285546d2b821cf421d4f5eb2576359d3770cd3
  • learning rate 1.0e-6, GRPO group size 4
  • seed schedule K=3: 101, 202, 303
  • the committed synthetic-scope certificate digest in docs/reward/proof_2026-09-01.json
  • ExtraDup mutants dup, extra, omit, unsubmit, claim
  • train patient ids p1..p8 vs holdout h1..h8
  • arms visual_only, certified_sor, shuffled, no-train
  • primary metric: holdout silent incorrect success
  • secondaries: honest-write success, over-halt, group-unscored rate

The freeze is so Phase-1/pilot training can't shop seeds. ExtraDup mutants stay out of the training reward. execute_seal stays false.

This PR doesn't report a training result.

Opened by an agent session, not the founder.

Pin checkpoint, lr, group size, K=3 seeds, certificate digest,
ExtraDup mutants, train/holdout patient ids, and arms before any
training step so the pilot cannot shop seeds.
@abrichr
abrichr merged commit 00fc435 into main Sep 2, 2026
2 checks passed
@abrichr
abrichr deleted the agent/certified-reward-m-freeze branch September 2, 2026 16:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant