anurag051194 commited on
Commit
fe33a25
·
verified ·
1 Parent(s): b95b830

Add full model card: results vs Granite 4.2 base and TwIL-LM3 peers, usage, GGUF, protocol and caveats

Browse files
Files changed (3) hide show
  1. .gitattributes +1 -0
  2. README.md +345 -80
  3. benchmarks.png +3 -0
.gitattributes CHANGED
@@ -38,3 +38,4 @@ Meridian-smaller-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
38
  Meridian-smaller-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
39
  Meridian-smaller-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
40
  Meridian-smaller-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
 
 
38
  Meridian-smaller-Q5_K_M.gguf filter=lfs diff=lfs merge=lfs -text
39
  Meridian-smaller-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
40
  Meridian-smaller-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
41
+ benchmarks.png filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -18,123 +18,388 @@ tags:
18
  - reinforcement-learning
19
  - grpo
20
  - gguf
 
21
  ---
22
 
23
  # Meridian-smaller
24
 
25
- Meridian-smaller is a 3.66B-parameter reasoning model built from
26
- [`ibm-granite/granite-4.2-3b`](https://huggingface.co/ibm-granite/granite-4.2-3b).
27
- It is specialized for formal-logic tasks through LoRA supervised fine-tuning, checkpoint
28
- fusion, WiSE-FT interpolation, and MGPO reinforcement learning.
29
 
30
- This repository contains the merged Transformers model and five llama.cpp GGUF builds.
 
 
 
 
 
31
 
32
- ## Model details
33
 
34
- - Architecture: Granite decoder-only dense transformer (`GraniteForCausalLM`)
35
- - Parameters: 3,659,737,600
36
- - Layers: 40
37
- - Hidden size: 2,560
38
- - Attention heads / KV heads: 40 / 8
39
- - Vocabulary size: 100,352
40
- - Checkpoint precision: bfloat16
41
- - Context window: 131,072 tokens inherited from Granite 4.2; release evaluations used shorter contexts
42
- - Published checkpoint: MGPO step 2580
43
- - MGPO schedule: resumed at step 800 with beta 0.02 and trained through step 2580
44
- - WiSE-FT base interpolation: 0.15
45
- - Reasoning format: emits a `<think>...</think>` block before the answer
46
 
47
- ## Usage with Transformers
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
48
 
49
- Use a recent Transformers release with Granite 4.2 support.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
50
 
51
  ```python
52
  import torch
53
  from transformers import AutoModelForCausalLM, AutoTokenizer
54
 
55
  model_id = "webAI-Official/Meridian-smaller"
56
- tokenizer = AutoTokenizer.from_pretrained(model_id)
57
  model = AutoModelForCausalLM.from_pretrained(
58
- model_id,
59
- dtype=torch.bfloat16,
60
- device_map="auto",
61
  )
62
 
63
- messages = [
64
- {
65
- "role": "user",
66
- "content": (
67
- "Does 'All dogs are mammals. Rex is a dog.' entail "
68
- "'Rex is a mammal'? Answer entailment, contradiction, or neutral."
69
- ),
70
- }
71
- ]
72
- inputs = tokenizer.apply_chat_template(
73
- messages,
74
- add_generation_prompt=True,
75
- tokenize=True,
76
- return_dict=True,
77
- return_tensors="pt",
78
  ).to(model.device)
79
 
80
- outputs = model.generate(**inputs, max_new_tokens=2048, do_sample=False)
81
- new_tokens = outputs[0, inputs["input_ids"].shape[-1]:]
82
- print(tokenizer.decode(new_tokens, skip_special_tokens=True))
83
  ```
84
 
85
- The packaged generation configuration enables sampling. Pass `do_sample=False` for greedy
86
- decoding. The chat template supports `enable_thinking=False` and
87
- `reasoning_effort="low"` through `chat_template_kwargs` when lower latency is preferred.
 
 
 
 
 
 
 
 
88
 
89
- ## GGUF / llama.cpp
90
 
91
- The following builds were produced with llama.cpp from the merged bfloat16 model:
 
 
92
 
93
- - `Meridian-smaller-Q4_K_M.gguf` — 2.09 GiB; recommended default
94
- - `Meridian-smaller-Q5_K_M.gguf` — 2.43 GiB
95
- - `Meridian-smaller-Q6_K.gguf` — 2.80 GiB
96
- - `Meridian-smaller-Q8_0.gguf` — 3.63 GiB
97
- - `Meridian-smaller-F16.gguf` — 6.82 GiB; reference and requantization build
 
 
98
 
99
  ```bash
100
- llama-cli -hf webAI-Official/Meridian-smaller:Q4_K_M \
101
- -cnv --temp 0 -n 2048
102
  ```
103
 
104
- F16 and Q8_0 were converted directly from the merged model. Q4_K_M, Q5_K_M, and Q6_K
105
- were quantized from the F16 GGUF without an importance matrix.
 
 
106
 
107
- ## Training
 
 
 
 
108
 
109
- The post-training pipeline had four stages:
 
 
110
 
111
- 1. LoRA supervised fine-tuning on synthetic formal-logic data.
112
- 2. Parameter-space fusion of selected SFT checkpoints.
113
- 3. WiSE-FT interpolation toward the Granite 4.2 base model, retaining 0.15 of the
114
- fine-tuned delta.
115
- 4. MGPO reinforcement learning with programmatic verifiers and partial-credit rewards.
116
 
117
- The MGPO run used group size 16, learning rate 5e-6, temperature 1.0, top-p 0.95,
118
- gamma 3.0, and beta 0.02 after the step-800 resume. This release is the merged step-2580
119
- policy, not a LoRA adapter.
120
 
121
- ## Evaluation note
 
 
 
 
 
 
 
 
 
 
 
 
122
 
123
- A checkpoint-selection probe used 100 prompts each from FOL translation, entailment, and
124
- math MCQ, with eight sampled completions per prompt (`temperature=0.8`, `top_p=0.95`).
125
- The macro Pass@1 was 0.4071 and macro Pass@8 was 0.5567. These are selection-probe
126
- metrics, not broad general-purpose benchmark claims, and were measured on the bfloat16
127
- Transformers weights rather than the GGUF builds.
128
 
129
- ## Limitations
 
 
 
 
 
 
 
 
 
 
130
 
131
- Meridian-smaller is specialized for formal logic and has not been comprehensively evaluated
132
- for general assistance, coding, multilingual use, safety, or very long contexts. It may produce
133
- incorrect reasoning or answers. It has no additional safety or preference tuning beyond what
134
- the base model carries. The published evaluation numbers do not measure the quantized files.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
135
 
136
  ## License and attribution
137
 
138
- Released under the webAI Non-Commercial License ver. 1.0; see `LICENSE.md`.
139
- The Granite 4.2 base model is Copyright IBM Corporation and is distributed under the
140
- Apache License 2.0, included as `apache-2.0-LICENSE.txt`.
 
 
 
 
 
 
18
  - reinforcement-learning
19
  - grpo
20
  - gguf
21
+ - meridian
22
  ---
23
 
24
  # Meridian-smaller
25
 
26
+ A 3.66B reasoning model for **formal logic** tasks, built from
27
+ [`ibm-granite/granite-4.2-3b`](https://huggingface.co/ibm-granite/granite-4.2-3b) through LoRA
28
+ supervised fine-tuning, checkpoint fusion, WiSE-FT weight interpolation, and entropy-weighted
29
+ GRPO reinforcement learning.
30
 
31
+ It improves in-domain formal-logic performance by **+28% relative** over its base model
32
+ (macro gate 0.431 → 0.554) **while holding held-out benchmark performance** — the 10-dataset
33
+ macro is level with the base (0.7942 → 0.7901) and the 14-dataset macro improves slightly
34
+ (0.7332 → 0.7425). On the same harness and sampled rows it reaches the highest Track A macro
35
+ gate, strict-7 and six-lane average of any arm in the tables below for which each can be
36
+ computed, including Qwen3-8B and gpt-oss-120b (the 120B has no gate or strict-7 value).
37
 
38
+ ![Meridian-smaller formal and general reasoning benchmarks against gpt-oss-120b, Qwen3-8B, LFM2.5-8B-A1B and TwIL-LM3](benchmarks.png)
39
 
40
+ ## Highlights
 
 
 
 
 
 
 
 
 
 
 
41
 
42
+ * **Large in-domain gain on the same harness.** Macro gate 0.4313 → 0.5539 (+0.123) and strict-7
43
+ 0.1821 → 0.2879 against its own base, measured on identical sampled rows. Both models are
44
+ heavily truncated at this budget (see [Limitations](#limitations-and-caveats)), and the base
45
+ more so, so the size of the gap is indicative rather than exact.
46
+ * **Top of the Track A summary rows at 3.66B.** Macro gate 0.5539 against Qwen3-8B's 0.5336
47
+ (2.2x the parameters) and TwIL-LM3's 0.4218; strict-7 0.2879 against 0.2093 and 0.1971;
48
+ six-lane average 0.5389 against gpt-oss-120b's 0.5192. The gate lead over Qwen3-8B is 0.020 —
49
+ smaller than the sampling noise at n = 200 per lane — so read that one as parity, not a win.
50
+ * **Strongest strict MCQ and language-model fit in the table.** `mcq_answer` strict accuracy
51
+ 0.4100 (next best 0.1200), `lean_critic` 0.7950 (tied with Qwen3-8B and its own base), and the
52
+ lowest `lm_corpus` and `math_corpus` perplexity of any comparable arm (2.3130 and 3.6983).
53
+ * **Holds general capability.** Track B 10-dataset macro 0.7901 and 14-dataset macro 0.7425 —
54
+ ahead of LFM2.5-8B-A1B (0.7884 / 0.7378) at less than half its parameter count, and behind
55
+ Qwen3-8B (0.8493 / 0.7591) and gpt-oss-120b (0.8689 / 0.8086). It gains on BBH-logic
56
+ (0.9013 → 0.9540), MATH-500 (0.6567 → 0.7467) and MuSR (0.5922 → 0.6409), and gives back
57
+ GSM-Symbolic (0.8900 → 0.8267) and ARC (0.8933 → 0.8600).
58
+ * **Structured formal output.** Tuned for the objects rather than the prose: FOL translation,
59
+ entailment labels, semantic parses, Lean statements and Lean proof critique.
60
+ * **Runs anywhere.** 3.66B parameters in bf16 (6.82 GiB), with a Q4\_K\_M GGUF at 2.09 GiB for
61
+ CPU or 4 GB of VRAM.
62
 
63
+ It is **not efficient**: it reasons at length. Track A generations average 1,902 tokens, against
64
+ 564 for TwIL-LM3, and 24.2% of them hit the length cap. It is also not a general assistant —
65
+ there is no safety or preference tuning here beyond what Granite 4.2 carries. See
66
+ [Limitations](#limitations-and-caveats).
67
+
68
+ ## Model Details
69
+
70
+ | Property | Value |
71
+ | ------------------------- | ---------------------------------------------------------------------------------------------- |
72
+ | Model ID | `webAI-Official/Meridian-smaller` |
73
+ | Base model | [`ibm-granite/granite-4.2-3b`](https://huggingface.co/ibm-granite/granite-4.2-3b) |
74
+ | Total parameters | 3.66B (3,659,737,600) |
75
+ | Architecture | Granite decoder-only dense transformer (`GraniteForCausalLM`); 40 layers, hidden size 2560, 40 attention heads / 8 KV heads |
76
+ | Input / output | Text / text |
77
+ | Language | English |
78
+ | Tokenizer vocabulary size | 100,352 |
79
+ | Context window | 131,072 tokens |
80
+ | Checkpoint precision | bfloat16 (6.82 GiB), plus Q4\_K\_M / Q5\_K\_M / Q6\_K / Q8\_0 / F16 GGUF builds |
81
+ | Post-training | LoRA SFT → checkpoint fusion → WiSE-FT (α = 0.15) → MGPO reinforcement learning (β = 0.02, step 2580) |
82
+ | Reasoning format | Emits a `<think>…</think>` block before the answer (default chat template) |
83
+ | Evaluated decoding | Greedy, 2048 new tokens (one retry at 4096), `max_seq_len` 8192 |
84
+ | Specialisation | Formal logic: FOL translation, entailment, semantic parsing, Lean formalisation and critique |
85
+ | License | webAI Non-Commercial License ver. 1.0 |
86
+
87
+ The base model's 131,072-token context is carried through unchanged, but every score on this card
88
+ was measured inside an 8,192-token window; longer contexts are inherited rather than validated
89
+ here.
90
+
91
+ ## Results
92
+
93
+ ### Track A — in-domain formal logic
94
+
95
+ Meridian-smaller and its base were run through the same harness, prompts, decoding settings and
96
+ sampled rows described under [Evaluation protocol](#evaluation-protocol), and the peer columns
97
+ are the values already published for TwIL-LM3 on its card, produced by that same harness (see
98
+ [Comparability](#limitations-and-caveats)). Throughput rows are reported where a dedicated
99
+ throughput measurement exists: `ans/s` is defined throughout as `tok/s ÷ mean generation length`,
100
+ so it measures completed answers rather than raw decode rate.
101
+
102
+ | lane / metric | Meridian-smaller | Granite-4.2-3B base | TwIL-LM3 | Llama-3.2-3B | LFM2-2.6B | LFM2.5-8B-A1B | Qwen3-8B | gpt-oss-120b ‡ |
103
+ |---|---:|---:|---:|---:|---:|---:|---:|---:|
104
+ | lean_formalize token_f1 | 0.5092 | 0.2943 | 0.5869 | 0.3690 | 0.1321 | 0.4655 | 0.4022 | **0.6306** |
105
+ | rule_induction derivation | 0.4195 | 0.2267 | 0.3192 | 0.0825 | 0.0615 | 0.1936 | 0.3680 | **0.6518** |
106
+ | entailment_label accuracy | 0.6700 | 0.3000 | 0.5750 | 0.3300 | 0.4700 | 0.5400 | 0.5800 | **0.7750** |
107
+ | mcq_answer accuracy | **0.4100** | 0.1200 | 0.1100 | 0.0000 | 0.0150 | 0.0750 | 0.0000 | 0.0700 |
108
+ | semantic_parse token_f1 | 0.4295 | 0.3910 | **0.4416** | 0.3102 | 0.3665 | 0.3778 | 0.4257 | 0.4331 |
109
+ | lean_critic accuracy | **0.7950** | **0.7950** | 0.6600 | 0.5300 | 0.5900 | 0.5500 | **0.7950** | 0.5550 |
110
+ | lm_corpus perplexity ↓ | **2.3130** | 2.6334 | 2.8972 | 2.8478 | 4.3815 | 4.9472 | 2.5440 | 912.23 § |
111
+ | math_corpus perplexity ↓ | **3.6983** | 4.4864 | 3.8229 | 4.7531 | 6.7472 | 8.3323 | 4.0083 | 1045.63 § |
112
+ | average, 6 lanes | **0.5389** | 0.3545 | 0.4488 | 0.2703 | 0.2725 | 0.3670 | 0.4285 | 0.5192 |
113
+ | **macro gate** | **0.5539** | 0.4313 | 0.4218 | 0.2925 | 0.3473 | 0.3757 | 0.5336 | — |
114
+ | **strict-7** | **0.2879** | 0.1821 | 0.1971 | 0.1229 | 0.1579 | 0.1714 | 0.2093 | — |
115
+ | macro_primary | **0.5875** | 0.4825 | 0.4475 | 0.3450 | 0.4188 | 0.4213 | 0.5750 | — |
116
+ | tok/s | — | — | 15880 | 16160 | **25230** | 22480 | 9420 | 3374 |
117
+ | mean gen length | 1902 | 2951 | **564** | 696 | 2296 | 1830 | 2094 | 1005 |
118
+ | **ans/s** | — | — | **28.1** | 23.2 | 10.9 | 12.0 | 4.5 | 3.4 |
119
+
120
+ ‡ **gpt-oss-120b** runs MXFP4 weights at tensor-parallel 2 — quantized and multi-GPU, so its
121
+ throughput rows are not directly comparable to the single-GPU BF16 arms. Its `procedural` lane
122
+ and the loose-match scorings were not collected, so the three summary rows below the six-lane
123
+ average cannot be computed for it; that is what the — cells mean, not a zero.
124
+
125
+ § The 120B's perplexities are three orders of magnitude off every other arm because its response
126
+ format and tokenizer make the corpus lanes score a different quantity. The number is reported
127
+ for completeness but is not a comparable measurement, and is excluded from the bolding.
128
+
129
+ **Throughput cells marked — for Meridian-smaller and its base.** The peer throughput columns come
130
+ from a separate dedicated throughput measurement that was not run for these two models, so
131
+ `tok/s` and `ans/s` are left blank rather than mixed with the in-run decode rate (which is on a
132
+ different basis). `mean gen length` is measured directly under the shared protocol and is
133
+ comparable across every column.
134
+
135
+ **`average, 6 lanes`** is the plain mean of the six objective rows above it, each at whatever
136
+ scoring that row reports (Lean F1, rule derivation, entailment, MCQ strict, semantic F1, critic).
137
+ It is a coarser summary than the three that follow — it mixes token-F1 with accuracy — but it is
138
+ the only summary row every arm here can be compared on, including the 120B.
139
+
140
+ The next three rows aggregate more carefully. None of them include the perplexity lanes or the
141
+ token-F1 scorings, which are not on a common 0–1 accuracy scale.
142
+
143
+ **`macro gate`** is the headline metric and the one the training pipeline gates on. It is the
144
+ equal-weight mean of five objectives: the four bounded classification lanes (`entailment_label`,
145
+ `mcq_answer`, `procedural`, `lean_critic`) plus `rule_induction`, scored by its continuous
146
+ derivation score. In the gate, `mcq_answer` and `procedural` are credited as
147
+ `max(exact_match, loose_match)`: for free-text answer lanes, a response that is correct but
148
+ differently formatted is a formatting artefact rather than a reasoning failure. This affects the
149
+ aggregate only — the per-lane rows above stay strict. (Meridian-smaller's loose-match MCQ is
150
+ 0.6500 against the 0.4100 strict figure shown in the lane row.)
151
+
152
+ **`macro_primary`** is the same mean over the four classification lanes alone, without
153
+ `rule_induction`.
154
+
155
+ **`strict-7`** is the mean of seven lanes scored under strict metrics only (`fol_translation`,
156
+ `entailment_label`, `mcq_answer`, `semantic_parse` and `lean_formalize` exact match,
157
+ `lean_critic` and `procedural` accuracy), with no loose-match credit anywhere. It is deliberately
158
+ harsh — exact match on generative lanes is near zero for every arm — so it is useful for ranking
159
+ models against each other but not as an absolute capability measure.
160
+
161
+ Meridian-smaller leads all four summary rows that every arm with a computable value can be
162
+ compared on. The clearest margins are over the arms at its own scale and above: 0.5539 against
163
+ 0.3757 on the gate for LFM2.5-8B-A1B, and 0.2879 against 0.1971 on strict-7 for TwIL-LM3. Against
164
+ Qwen3-8B the gate gap is only 0.020, but strict-7 is 0.2879 against 0.2093 — a difference that
165
+ does not depend on loose-match credit — and Meridian-smaller leads on MCQ (strict 0.4100 against
166
+ 0.0000), rule induction (0.4195 against 0.3680) and both perplexity lanes.
167
+
168
+ It does not lead every lane. gpt-oss-120b is clearly stronger on `rule_induction` (0.6518),
169
+ entailment (0.7750) and Lean formalisation (0.6306), and TwIL-LM3 remains ahead on Lean
170
+ formalisation (0.5869 against 0.5092) and semantic parsing (0.4416 against 0.4295). The two weak
171
+ spots in absolute terms are `procedural` (strict 0.1200, loose 0.2350) and FOL translation
172
+ (exact match 0.0100).
173
+
174
+ ### Track B — held-out benchmarks
175
+
176
+ | dataset | Meridian-smaller | Granite-4.2-3B base | TwIL-LM3 | Llama-3.2-3B | LFM2-2.6B | LFM2.5-8B-A1B | Qwen3-8B | gpt-oss-120b ‡ |
177
+ |---|---:|---:|---:|---:|---:|---:|---:|---:|
178
+ | gsm8k | 0.9433 | 0.9533 | 0.8733 | 0.8300 | 0.8767 | 0.9133 | 0.9567 | **0.9767** |
179
+ | svamp | **0.9500** | 0.9200 | 0.8500 | 0.8200 | 0.9000 | 0.9133 | 0.9400 | 0.9400 |
180
+ | gsm_symbolic | 0.8267 | 0.8900 | 0.7567 | 0.8067 | **0.9767** | 0.9267 | 0.8133 | 0.8467 |
181
+ | arc_cot | 0.8600 | 0.8933 | 0.8467 | 0.7967 | 0.8667 | 0.9033 | 0.9633 | **0.9667** |
182
+ | logicbench | 0.7933 | 0.7733 | 0.7167 | 0.5733 | 0.6267 | 0.7200 | **0.8567** | 0.8533 |
183
+ | strategyqa | 0.5967 | 0.6267 | 0.6500 | 0.6533 | 0.6433 | 0.6667 | 0.7400 | **0.7867** |
184
+ | drop | 0.7367 | 0.7467 | 0.7467 | 0.6733 | 0.6900 | 0.6633 | **0.8833** | 0.8500 |
185
+ | csqa | 0.7667 | 0.7733 | 0.7367 | 0.7500 | 0.7433 | 0.7700 | **0.8633** | 0.8367 |
186
+ | musr | 0.6409 | 0.5922 | 0.4957 | 0.4932 | 0.4867 | 0.5703 | 0.6301 | **0.6852** |
187
+ | mmlu_redux | 0.7867 | 0.7733 | 0.6667 | 0.6000 | 0.7133 | 0.8367 | 0.8500 | **0.9467** |
188
+ | ifeval | 0.7500 | 0.7633 | 0.6433 | 0.7167 | 0.7300 | **0.8900** | 0.8400 | 0.7900 |
189
+ | rudas_ood | 0.0437 ¶¶ | 0.0012 | 0.0365 | **0.0733** | 0.0017 | 0.0061 | 0.0468 | 0.0000 ¶ |
190
+ | bbh_logic | 0.9540 | 0.9013 | 0.6633 | 0.5333 | 0.5713 | 0.7700 | 0.6367 | **0.9980** |
191
+ | math500 | 0.7467 | 0.6567 | 0.6900 | 0.4233 | 0.7133 | 0.7800 | 0.6100 | **0.8433** |
192
+ | **macro (10 CoT datasets)** | 0.7901 | 0.7942 | 0.7339 | 0.6997 | 0.7523 | 0.7884 | 0.8493 | **0.8689** |
193
+ | **macro (all 14)** | 0.7425 | 0.7332 | 0.6694 | 0.6245 | 0.6814 | 0.7378 | 0.7591 | **0.8086** |
194
+ | tok/s | — | — | 15880 | 16160 | **25230** | 22480 | 9420 | 3374 |
195
+ | mean gen length | ≈792 | ≈1282 | **482** | 510 | ≈796 | ≈1327 | ≈1931 | 801 |
196
+ | **ans/s** | — | — | **32.9** | 31.7 | ≈31.7 | ≈16.9 | 4.9 | 4.2 |
197
+
198
+ ‡ MXFP4 weights, tensor-parallel 2 — quantized and multi-GPU, so not directly comparable to the
199
+ single-GPU BF16 rows. ¶ 74% of its `rudas_ood` generations hit the length cap, so that cell is a
200
+ truncation artefact rather than a measured score; excluding the row, its 13-dataset macro is
201
+ 0.8708. ¶¶ 93.7% of Meridian-smaller's `rudas_ood` generations also hit the length cap, so its
202
+ cell is likewise a truncation artefact rather than a measurement of the model's ability.
203
+
204
+ Lengths marked ≈ are derived from stored generations rather than read from the run. For
205
+ Meridian-smaller and its base they were re-tokenized directly with the model's own tokenizer and
206
+ averaged over the 14 datasets (MuSR counted once); the same method reproduces TwIL-LM3's measured
207
+ 482 to within 1%. The peer lengths marked ≈ use each model's characters-per-token ratio and are
208
+ carried over from the TwIL-LM3 card. `tok/s` and `ans/s` are blank for Meridian-smaller and its
209
+ base for the reason given under the Track A table.
210
+
211
+ The honest summary of this table is that Meridian-smaller does not lead it. Larger models score
212
+ higher, and gpt-oss-120b leads seven of the fourteen dataset rows. Three things are worth
213
+ extracting anyway. First, it holds its own base on the held-out suite (10-dataset macro 0.7901
214
+ against 0.7942, a difference well inside the sampling noise at n = 300 per dataset) while gaining
215
+ in-domain, which is the point of the WiSE-FT stage. Second, the 14-dataset macro rises 0.0093
216
+ over the base, driven by BBH-logic, MATH-500 and MuSR. Third, it edges LFM2.5-8B-A1B on both
217
+ macros at less than half the parameters and leads the table outright on SVAMP (0.9500).
218
+
219
+ Track B here was run with the chat template's thinking mode **disabled** for Meridian-smaller and
220
+ its base (the prompt ends in an empty `<think></think>`), as it was for TwIL-LM3, whereas Track A
221
+ uses the default thinking mode. The Track B numbers therefore describe non-reasoning behaviour;
222
+ they are not a measure of what a thinking-mode generation would score. Some Track B cells are also
223
+ truncation-limited: MATH-500 hits the cap on 14.0% of rows, and SVAMP, GSM-Symbolic and
224
+ MuSR-team sit slightly above the 2% cap-hit threshold (3.0%, 3.0% and 2.8%).
225
+
226
+ ### Checkpoint selection
227
+
228
+ MGPO checkpoints were compared on both tracks, and step 2580 was published because it has the
229
+ best (or tied-best) held-out score rather than the best in-domain gate:
230
+
231
+ | checkpoint | macro gate | macro_primary | B10 | B14 | Track A truncation |
232
+ |---|---:|---:|---:|---:|---:|
233
+ | WiSE-FT α = 0.15 (RL initialiser) | 0.531 | 0.583 | — | — | 27.6% |
234
+ | MGPO step 1200 | **0.569** | **0.598** | 0.7819 | 0.7362 | 21.0% |
235
+ | MGPO step 2000 | 0.557 | 0.581 | 0.7882 | 0.7425 | 23.1% |
236
+ | MGPO step 2200 | 0.567 | 0.583 | 0.7865 | 0.7406 | 23.3% |
237
+ | **MGPO step 2580 (published)** | 0.554 | 0.588 | **0.7901** | **0.7425** | 24.2% |
238
+
239
+ Track A gate here uses the same definition as the tables above (rule induction included).
240
+ Step 1200 leads the gate by 0.015 and `macro_primary` by 0.010, but step 2580 has the highest
241
+ 10-dataset macro and ties step 2000 on the 14-dataset macro (0.7425 for both at four decimals),
242
+ and the differences between the later checkpoints on Track A are within sampling noise at
243
+ n = 200. The checkpoint-selection probe recorded for this release (100 prompts each from FOL
244
+ translation, entailment and math MCQ, eight sampled completions per prompt at
245
+ `temperature = 0.8`, `top_p = 0.95`) gave macro Pass@1 0.4071 and macro Pass@8 0.5567. That probe
246
+ is a selection tool, not a benchmark claim.
247
+
248
+ ## Usage
249
 
250
  ```python
251
  import torch
252
  from transformers import AutoModelForCausalLM, AutoTokenizer
253
 
254
  model_id = "webAI-Official/Meridian-smaller"
255
+ tok = AutoTokenizer.from_pretrained(model_id)
256
  model = AutoModelForCausalLM.from_pretrained(
257
+ model_id, torch_dtype=torch.bfloat16, device_map="auto"
 
 
258
  )
259
 
260
+ messages = [{"role": "user", "content":
261
+ "Does 'All dogs are mammals. Rex is a dog.' entail 'Rex is a mammal'? "
262
+ "Answer entailment, contradiction, or neutral."}]
263
+ inputs = tok.apply_chat_template(
264
+ messages, add_generation_prompt=True,
265
+ return_tensors="pt", return_dict=True,
 
 
 
 
 
 
 
 
 
266
  ).to(model.device)
267
 
268
+ out = model.generate(**inputs, max_new_tokens=2048, do_sample=False)
269
+ print(tok.decode(out[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))
 
270
  ```
271
 
272
+ `return_dict=True` matters on transformers 5.x, where `apply_chat_template` returns a
273
+ `BatchEncoding` rather than a bare tensor; the above works on both 4.x and 5.x. Use a
274
+ Transformers release with Granite 4.2 support.
275
+
276
+ The reported numbers use **greedy decoding** (`do_sample=False`) and a **2048-token** generation
277
+ budget. Note that the shipped `generation_config.json` enables sampling (`do_sample=true`,
278
+ `temperature=1.0`, `top_p=0.95`), so `do_sample=False` must be passed explicitly to reproduce the
279
+ evaluation. By default the chat template opens a `<think>` block, so the model reasons before it
280
+ answers and a short generation budget truncates that reasoning and scores far worse. The template
281
+ also accepts `enable_thinking=False` (empty think block, lower latency and lower quality on
282
+ reasoning-heavy tasks) and `reasoning_effort="low"` through `chat_template_kwargs`.
283
 
284
+ ### GGUF / llama.cpp
285
 
286
+ Quantized GGUF builds ship in this repository alongside the safetensors weights. The Granite
287
+ architecture is supported by llama.cpp; the model uses a ChatML-style template with `<|im_end|>`
288
+ as EOS, so run it in conversation mode (`-cnv`).
289
 
290
+ | file | quant | size | bits/weight | notes |
291
+ |---|---|---:|---:|---|
292
+ | `Meridian-smaller-Q4_K_M.gguf` | Q4_K_M | 2.09 GiB | 4.91 | recommended default; runs on CPU or 4 GB of VRAM |
293
+ | `Meridian-smaller-Q5_K_M.gguf` | Q5_K_M | 2.43 GiB | 5.71 | a little more headroom than Q4_K_M |
294
+ | `Meridian-smaller-Q6_K.gguf` | Q6_K | 2.80 GiB | 6.57 | close to Q8_0 quality at about three-quarters the size |
295
+ | `Meridian-smaller-Q8_0.gguf` | Q8_0 | 3.63 GiB | 8.51 | near-lossless, for quality-sensitive use |
296
+ | `Meridian-smaller-F16.gguf` | F16 | 6.82 GiB | 16.01 | unquantized, for requantization or reference runs |
297
 
298
  ```bash
299
+ llama-cli -hf webAI-Official/Meridian-smaller:Q4_K_M -cnv --temp 0 -n 2048
 
300
  ```
301
 
302
+ Two things matter for reproducing the scores above under llama.cpp. Pass `--temp 0`, because the
303
+ evaluation is greedy while the packaged sampling defaults are not. And leave the generation
304
+ budget large — 2048 tokens or more — since the model emits a `<think>` block before answering
305
+ and a short budget truncates it, which costs far more accuracy than the quantization does.
306
 
307
+ F16 and Q8_0 were produced directly by `convert_hf_to_gguf.py` from the released bf16 weights; the
308
+ K-quants (Q4_K_M, Q5_K_M, Q6_K) were quantized from the F16 build with `llama-quantize`, without
309
+ an importance matrix. F16 is not bit-identical to the released weights: bf16 and f16 carry the
310
+ same 16 bits but trade exponent range against mantissa precision, so the conversion is a
311
+ narrowing one, in practice negligible for inference.
312
 
313
+ The published Track A and Track B numbers were measured on the **bf16** weights through vLLM, not
314
+ on any of these GGUF builds, so expect small deviations — most likely at Q4_K_M — that have not
315
+ been quantified here.
316
 
317
+ ## How it was built
 
 
 
 
318
 
319
+ Four stages on top of the base model:
 
 
320
 
321
+ 1. **LoRA supervised fine-tuning** on a synthetic formal-logic corpus covering the Track A
322
+ objectives (first-order-logic translation, entailment labelling, semantic parsing, Lean
323
+ formalisation and critique, procedural reasoning, rule induction).
324
+ 2. **Checkpoint fusion** — parameter-space averaging of intermediate SFT checkpoints, rather
325
+ than taking the final checkpoint.
326
+ 3. **WiSE-FT interpolation** toward the pretrained base, `W = (1 − α)·W_base + α·W_finetuned`
327
+ with **α = 0.15** — i.e. only 15% of the fine-tuned delta is retained. This conservative
328
+ interpolation is the direct reason held-out capability survives.
329
+ 4. **MGPO** — entropy-weighted GRPO reinforcement learning against a programmatic verifier, with
330
+ partial credit for loose matches and token-F1 so that all-fail prompt groups still produce
331
+ gradient. Group size 16, learning rate 5e-6, sampling temperature 1.0, top-p 0.95,
332
+ γ = 3.0. The run was resumed at step 800 with β = 0.02 and trained through step 2580, and the
333
+ published checkpoint is **step 2580**. This release is the merged policy, not a LoRA adapter.
334
 
335
+ ## Limitations and caveats
 
 
 
 
336
 
337
+ **Truncation.** At a 2048-token budget (one retry at 4096), **24.2%** of Track A generations
338
+ hit the cap. That is far above the 2% threshold our protocol requires to mark a comparison
339
+ `rankable`, so the Track A numbers are **not rankable** and should be read as indicative rather
340
+ than exact. The base is worse (49.2%), and both are pessimistic because a truncated response
341
+ scores zero regardless of reasoning quality — so the true Track A gap over the base is probably
342
+ narrower than +0.123, and part of the improvement is shorter generations rather than better
343
+ answers. For context, Qwen3-8B truncates 23.9% of rows, LFM2.5-8B-A1B 17.9%, LFM2-2.6B 41.3% and
344
+ Llama-3.2-3B 10.0% under the same budget, TwIL-LM3 4.4%. Some Track B stages are also flagged
345
+ unrankable for the same reason (`rudas_ood` 93.7% cap-hit, `math500` 14.0%, plus marginal excess
346
+ on SVAMP, GSM-Symbolic and MuSR-team); the held-out core, retention, IFEval and log-likelihood
347
+ stages are rankable.
348
 
349
+ **Verbose by construction.** Track A generations average 1,902 tokens and Track B generations
350
+ about 792, so cost per answer is substantially higher than the TwIL-LM family (564 and 482 tokens)
351
+ even though quality per answer is higher on Track A.
352
+
353
+ **Scope.** Tuned for formal logic. The Track B suite does not cover code generation or tool use
354
+ (HumanEval, LiveCodeBench and BFCL were not run for this model or its base), so this release
355
+ makes no claim about those. The weak absolute areas inside the specialisation are FOL translation
356
+ (exact match 0.0100), `procedural` (strict 0.1200) and semantic parsing exact match (0.0000);
357
+ `rule_induction` parses only 56.5% of outputs.
358
+
359
+ **Not a chat model.** It was optimised against automatic verifiers on logic tasks. It has had no
360
+ safety tuning beyond whatever the base model carries, and no instruction-following alignment
361
+ work — IFEval is 0.7500 against the base's 0.7633.
362
+
363
+ **Comparability.** For Track A, Meridian-smaller, its base, Qwen3-8B, LFM2-2.6B, LFM2.5-8B-A1B
364
+ and Llama-3.2-3B were checked to share the same sampled-row manifest and dataset hash, seed and
365
+ decoding; the TwIL-LM3 and gpt-oss-120b values are carried over from the TwIL-LM3 card, which
366
+ describes the same harness. For Track B, the arms checked share the same sampled rows and
367
+ decoding, but the serving engine differs between arms (vLLM 0.19.1 for Meridian-smaller, its base,
368
+ Qwen3-8B and LFM2.5-8B-A1B; vLLM 0.11.2 for TwIL-LM3 and Llama-3.2-3B), and the engine version is
369
+ part of the protocol hash. With n = 200 per lane on Track A and n = 300 per dataset on Track B,
370
+ differences of two to three points are within sampling noise.
371
+
372
+ ## Evaluation protocol
373
+
374
+ - Track A: `n = 200` per objective, greedy (`temperature = 0`), `max_new_tokens = 2048`, one
375
+ retry at 4096 for truncated rows, `max_seq_len = 8192`, seed 42, default (thinking-enabled)
376
+ chat template.
377
+ - Track B: 300 examples per task, greedy, `max_gen_toks = 4096`, `max_model_len = 8192`,
378
+ `repetition_penalty = 1.0`, chat template applied with thinking disabled, vLLM backend.
379
+ - Both tracks use the same protocol for the model and its base, in a paired run over identical
380
+ sampled rows.
381
+
382
+ `repetition_penalty = 1.0` is load-bearing. A 1.1 penalty produced apparent 20-point swings on
383
+ Track B that were pure decoding artefact; the decoding kwargs are hashed into the protocol
384
+ identity so a mismatched runner fails loudly instead of quietly producing a different number.
385
+
386
+ ## Relationship to TwIL-LM
387
+
388
+ Meridian-smaller applies the same post-training pipeline as the
389
+ [TwIL-LM3](https://huggingface.co/webAI-Official/TwIL-LM3) and TwIL-LM family — LoRA SFT,
390
+ checkpoint fusion, WiSE-FT and MGPO — to a different base, IBM's Granite 4.2 3B, instead of
391
+ SmolLM3 or SmolLM2. Compared with TwIL-LM3 it is a stronger in-domain model (macro gate 0.5539
392
+ against 0.4218) and a stronger held-out one (10-dataset macro 0.7901 against 0.7339), at the
393
+ price of much longer generations and a much higher truncation rate. Like the TwIL-LM models, it
394
+ ships as a full merged model on `main`, loaded directly with `AutoModelForCausalLM`.
395
 
396
  ## License and attribution
397
 
398
+ Released under the **webAI Non-Commercial License ver. 1.0** — see `LICENSE.md` in this
399
+ repository.
400
+
401
+ The base model, [`ibm-granite/granite-4.2-3b`](https://huggingface.co/ibm-granite/granite-4.2-3b),
402
+ is Copyright IBM Corporation and is distributed under the Apache License 2.0; its licence text is
403
+ retained as `apache-2.0-LICENSE.txt` and all credit for the base model goes to IBM. Apache 2.0
404
+ permits distributing derivative works under different terms provided attribution is preserved,
405
+ which is what the pair of licence files in this repository does.
benchmarks.png ADDED

Git LFS Details

  • SHA256: a940575cbd32424f6b2a3454746aaa286162e5b24155b4470f01c8bf2d9e6a92
  • Pointer size: 131 Bytes
  • Size of remote file: 139 kB