Raiff1982 commited on
Commit
016b903
·
verified ·
1 Parent(s): c7fa0ed

Final result: 28.0% — complete STaR did not beat baseline

Browse files
Files changed (1) hide show
  1. README.md +33 -46
README.md CHANGED
@@ -1,58 +1,45 @@
1
  ---
 
2
  base_model: Raiff1982/codette-llama-3.1-8b-merged
3
- library_name: transformers
4
- model_name: codette-newton-star-r
5
- tags:
6
- - generated_from_trainer
7
- - sft
8
- - trl
9
- licence: license
10
  ---
11
 
12
- # Model Card for codette-newton-star-r
13
 
14
- This model is a fine-tuned version of [Raiff1982/codette-llama-3.1-8b-merged](https://huggingface.co/Raiff1982/codette-llama-3.1-8b-merged).
15
- It has been trained using [TRL](https://github.com/huggingface/trl).
 
 
 
 
16
 
17
- ## Quick start
 
 
 
18
 
19
- ```python
20
- from transformers import pipeline
 
21
 
22
- question = "If you had a time machine, but could only go to the past or the future once and never return, which would you choose and why?"
23
- generator = pipeline("text-generation", model="Raiff1982/codette-newton-star-r", device="cuda")
24
- output = generator([{"role": "user", "content": question}], max_new_tokens=128, return_full_text=False)[0]
25
- print(output["generated_text"])
26
- ```
27
 
28
- ## Training procedure
29
 
30
-
 
 
 
 
 
31
 
 
 
 
 
 
 
32
 
33
-
34
- This model was trained with SFT.
35
-
36
- ### Framework versions
37
-
38
- - TRL: 1.8.0
39
- - Transformers: 4.57.6
40
- - Pytorch: 2.10.0+cu128
41
- - Datasets: 5.0.0
42
- - Tokenizers: 0.22.2
43
-
44
- ## Citations
45
-
46
-
47
-
48
- Cite TRL as:
49
-
50
- ```bibtex
51
- @software{vonwerra2020trl,
52
- title = {{TRL: Transformers Reinforcement Learning}},
53
- author = {von Werra, Leandro and Belkada, Younes and Tunstall, Lewis and Beeching, Edward and Thrush, Tristan and Lambert, Nathan and Huang, Shengyi and Rasul, Kashif and Gallouédec, Quentin},
54
- license = {Apache-2.0},
55
- url = {https://github.com/huggingface/trl},
56
- year = {2020}
57
- }
58
- ```
 
1
  ---
2
+ license: llama3.1
3
  base_model: Raiff1982/codette-llama-3.1-8b-merged
4
+ library_name: peft
5
+ tags: [codette, star, rationalization, lora, negative-result, research-artifact]
 
 
 
 
 
6
  ---
7
 
8
+ # newton-star-r — Complete STaR: keep-correct + rationalization (final result: 28.0%)
9
 
10
+ Fourth and final arm of a controlled STaR study — the first to implement the
11
+ **complete** method from Zelikman et al.: 350 keep-correct chains
12
+ (difficulty-matched MMLU-Pro STEM) **plus 180 rationalized chains** built from
13
+ problems the model originally got *wrong* (correct answer supplied during
14
+ generation, model derives the reasoning, hint stripped from the training
15
+ example, anti-leak filter rejecting chains that cite being given the answer).
16
 
17
+ **Result: 28.0% on GPQA-main (reason mode, n=100).** Rationalization recovered
18
+ the easy-arm regression (25.0% -> 28.0%) but did **not** exceed difficulty-matched
19
+ keep-correct (also 28.0%) or the 34.0% untrained baseline. The widely-held
20
+ assumption that rationalization closes the keep-correct gap did not hold here.
21
 
22
+ Two measured factors bound its contribution at 8B: ~9% of failures were
23
+ unconstructible even with the correct answer given, and answer-scaffolded
24
+ chains may encode the conclusion without the search a cold solve requires.
25
 
26
+ We publish this exactly as measured — a benchmark that can't be trusted to
27
+ report failure can't be trusted to report success.
 
 
 
28
 
29
+ ## The STaR Study (GPQA-main, reason mode, n=100 per arm)
30
 
31
+ | Adapter | Training data | GPQA | Verdict |
32
+ |---|---|---|---|
33
+ | newton (untrained baseline) | — | **34.0%** | reproduced to the decimal, 4 days apart |
34
+ | newton-star | 500 easy-science keep-correct | **25.0%** | regressed to chance |
35
+ | newton-star-hard | 350 MMLU-Pro STEM keep-correct | **28.0%** | attenuated the harm, below baseline |
36
+ | newton-star-r | 350 keep-correct + 180 rationalized | **28.0%** | complete method; recovered easy-arm damage, still below baseline |
37
 
38
+ **Finding:** neither half of STaR — keep-correct nor rationalization, nor both
39
+ together — beat the untrained baseline at 8B scale. Rationalization recovered
40
+ the easy-arm regression (25.0% -> 28.0%) but did not exceed keep-correct-hard
41
+ or the 34.0% baseline. Self-taught reasoning *consolidates existing ability
42
+ rather than extending it.* Full methodology, controls, and changelogs:
43
+ [Codette-Reasoning](https://github.com/Raiff1982/Codette-Reasoning).
44
 
45
+ Created by Jonathan Harrison (Raiff1982) · Raiff's Bits LLC