Model X-Ray evidence correction β€” 6 September 2026. Any structural-location, knowledge-separation, portrait-visualization, lesion-response, legacy simulated-quantization, robustness or prior X-Ray endorsement previously linked from this card has been withdrawn. It is not current evidence. The withdrawal remains in force; the separately dated remeasurement below does not reinstate those claims. Correction record.

VG1 status Β· 22 September 2026 (Europe/Istanbul)

The 7B class opened on 21 September 2026. Any registered account may scan the repositories on the eligible list of the scope page β€” exact Apache-2.0 revisions, listed there with their status β€” within the free allowance of 20 browser scans and five distinct source models per calendar month. 7B-class repositories run as single-model quantization simulations. A comparison of two 7B-class checkpoints is currently accepted by the submission form but fails inside the worker sandbox (under investigation, 22 September 2026); please do not submit one until this sentence is removed (a job that fails this way is recorded as failed and the scan it reserved is returned automatically). HuggingFaceTB/SmolLM3-3B-Base and HuggingFaceTB/SmolLM3-3B are listed on the register, but the current instrument release does not support their tokenizer contract, so those submissions fail as well. Scans run on Tetracta-operated local GPU workers; no cloud GPU is used in this beta. Models above the 7B class are not scanned in this beta. This card's presence is not an eligibility grant. External customer acceptance remains pending; results are research-stage and payments are disabled.

A report describes its recorded artifacts and measurement conditions; it is not a model-quality ranking, safety certificate or deployment verdict. A quantization simulation does not create a deployable quantized model. Knowledge scans remain unavailable and earlier withdrawn claims remain withdrawn. Hallucination, undertraining and overtraining remain research questions, not measured product features.

Service status Β· Current scope and report guide Β· Correction record Β· Report examples

Model X-Ray status Β· 12 September 2026 (Europe/Istanbul) (12 September 2026 β€” superseded by the 22 September entry above; "Customer jobs run on RunPod" no longer applies.)

VG1 has been deployed. Customer scan-to-report acceptance is still pending.

The free first-beta scope covers eligible public models up to and including 7B, within each account's allowance, for checkpoint comparisons and in-memory quantization simulations. Customer jobs run on RunPod. The service selector determines supported models, revisions and scan types; this card's presence is not an eligibility grant. Later Pro access above 7B requires an explicit grant and the enabled account limits.

Knowledge scans remain unavailable. Reports describe recorded artifacts and measurement conditions, not model-quality rankings, safety certificates or deployment verdicts. A quantization simulation does not produce a deployable quantized model. Hallucination, undertraining and overtraining remain research questions, not measured product features. Earlier withdrawn X-Ray interpretations remain withdrawn; the historical study and model measurements below are not new VG1 customer results.

Service status Β· Current scope and report guide Β· Dated correction

New current-VG1 Twin remeasurement Β· 12 September 2026

Five published step-30000 checkpoints were scanned again on 12 September 2026. The results below come from new research measurements; the July measurements remain historical. The original local checkpoint tensors were verified exactly against the published safetensors artifacts at revision cf7cf00fbe8f99cf49a549ecf831fc45cae66dbb.

Seven actual GPU scans were completed: the original Vanilla and Rational checkpoints were each loaded and measured twice in separate processes, and the three published seed-control checkpoints were each measured once. Ten comparisons reuse those seven observations; two of the ten are algebraic self-checks using an observation against itself, not independent scans.

New recorded check Bounded result
Two separate-process, same-checkpoint repeat comparisons No response finding; greedily generated token sequences matched.
Original cross-family comparison and the second cross-family seed pair Response differences and different greedily generated token sequences were observed.
Four within-family seed comparisons Response differences and different greedily generated token sequences were also observed.
Two algebraic self-checks Zero finding, as required; not additional experimental runs.

This does not establish an architecture effect, seed independence or a quality ranking. Differences also occur between seeds within the same family. The original cross-family pair alone cannot identify what caused its difference. No model-specific statistical noise envelope or population conclusion is certified by this small study. The historical retractions below remain in force.

New Fable reports, artifact links and controls Β· Pinned HTML archive. The reports distinguish real measurement dates from historical checkpoint dates and identify the exact public artifact paths. They are unsigned research copies, not customer acceptance records.

The new execution uses the current VG1 numerical instrument and Fable presentation through an explicit private custom-model integration. Native-model and adapter outputs were checked for equality before each scan, and the implementation boundary is recorded separately. The ordinary service still does not accept these archived research checkpoints; this does not add custom-operator support to the customer selector. Proprietary implementations, tokenizer, configuration details and raw records are not published. No new quantization-simulation or knowledge result is claimed by this remeasurement.

πŸ“š Current scope and report guide: https://www.tetracta.ai/model-xray/scope/

Current availability: see the dated VG1 status above. Customer scan-to-report acceptance remains pending.

Twin Loss Curves β€” and the Seed-Null That Retracted the Headline

Correction (13 July 2026). The first version of this card reported that the two models below compute measurably differently inside β€” a "4.4Γ— quieter internals under int8" figure, an "84% of depth" divergence, and a routing-stability edge. We then ran the one control we had flagged as "obvious next": an independent-seed null. We trained three more models and re-scanned. Every one of those differential-internal claims fell inside the noise between two ordinary runs, and is retracted. This card now documents the original claims, the control that refuted them, and what survived. The retraction is the point β€” see the full post-mortem at https://www.tetracta.ai/note-xray-twin-study.html.

Historical update (17 July 2026; superseded). The old note described a re-run under rs-1.3 and later referred to rs-1.5 / ps-1.1 / mv-1.2. These are historical version labels, not the current VG1 pipeline. It reported byte-identical int8 outputs (0/6 behavior change on either twin) alongside different internal profiles. Those old profiles are not current evidence of architecture effects, valid internal resolution or added value over output comparisons. The seed-null retraction and the 6 September withdrawal remain in force. Current report contract: rs-1.7 / mv-1.4, with rp-1.3 presentation. The new 12 September private research results are reported separately above; the historical results below have not been relabeled. Current availability: https://www.tetracta.ai/xray.html

The original training record describes two 0.93B language models with matched recorded conditions β€” the same seed, data order and schedule β€” and an intended change in the attention weighting function (standard softmax versus Tetracta's undisclosed rational operator, zero extra params). These training records have not been independently revalidated by the current VG1 update.

vanilla/ rational/
val bits-per-byte @ step 30 000 1.0092 1.0102
parameters 0.929 B 0.929 B

The recorded validation-loss values are close. This limited observation does not establish equality on every output-level task or prove architecture equivalence.

What we first claimed, and what the control did to it

The historical study scanned both with the then-used Tetracta Model X-Ray development pipeline and published four internal findings. A reviewer asked for the denominator: how much do two runs of the same recipe, differing only in seed, already diverge inside? We trained VAN-seed-B, VAN-seed-C (two more softmax runs) and RAT-seed-B, and measured, with thresholds sealed in advance. Five checkpoints; ~$300 compute; all pods terminated; all checkpoints md5-sealed.

Historical training-loss observations β€” not a new VG1 result. The three softmax seeds landed at bpb 1.0092 / 1.0072 / 1.0085; the two rational seeds at 1.0102 / 1.0101. The spread between softmax seeds (0.0020 bpb) is larger than the original softmax-vs-rational gap (0.0010). For these reported checkpoints, the original between-arm loss gap is smaller than the observed between-seed spread. That is not an equivalence test or a general seed-independent result.

"4.4Γ— quieter inside" β€” RETRACTED. Internal disturbance under int8, per model: VAN-42 4.166, VAN-B 0.573, VAN-C 0.967, RAT-42 1.277, RAT-B 1.104. Spread among softmax seeds alone: 7.3Γ—. The original VAN-42 was a high outlier; the softmax median (0.97) sits below rational (1.19). The direction did not survive, let alone the 4.4Γ—.

Routing-stability edge β€” RETRACTED. int8 token-flip rate: VAN-42 2.97%, VAN-B 0.93%, VAN-C 1.74%, RAT-42 0.52%, RAT-B 2.66%. The rational replication lands inside the softmax band; RAT-42's low value was a lucky draw.

"84% of depth" divergence β€” RETRACTED (as a ceiling). Cross-scan mean deviation: operator pair 0.000307, seed pair 0.000324 β€” both span 84% of depth (ratio 1.06Γ—). Two ordinary softmax runs diverge internally exactly as much as the operator swap does. Our own note had flagged 84% as a ceiling indicator; the control confirmed it, at our expense.

A null we kept (unchanged). An early read suggested rational abstains more on trick questions. Adversarial re-verification killed it on day one (a generation-budget artifact, pβ‰ˆ0.2). Still a null.

Other historical claims β€” not current VG1 evidence

The earlier card reported a separate Qwen1.5-MoE-A2.7B experiment involving 1.47M routing decisions and approximately 2.6% versus 1.3% routing changes for a real int8 kernel and a simulation. The claim that this establishes a general, seed-independent twofold kernel/simulation effect is not supported by current VG1 evidence and must not be carried forward. No new reproduction of that experiment is published in this update.

A legacy note also reported a bespoke integration with the undisclosed operator. This is not a supported-current-product claim. The ordinary current scanner does not accept this repository's archived layout, and the rational arm cannot be evaluated correctly without its private implementation. The separately validated 12 September private integration establishes only the bounded research executions described above; it does not change ordinary service eligibility.

Honest limits

Historical setup: 0.93B models at step 30k of 157k β€” young. Simulated per-row RTN quantization, not production GPTQ/AWQ. Behavioural comparisons rest on n=6 probe prompts. Nothing here is a scaling law, and nothing here claims rational produces a better model. The internal-difference claims we did make were retracted by the original control; the later withdrawal also applies. These records do not validate the current release.

Files

vanilla/  model-step{10000,20000,30000}.safetensors + config.json   # softmax baseline
rational/ model-step{10000,20000,30000}.safetensors + config.json   # Tetracta rational attention
xray/     portrait-{vanilla,rational}-step30000.png
xray_summary.json    weight_error.json + reproduce_weight_error.py    manifest.json (sha256)

float32. lm_head is tied to tok_emb. rational/ cannot be run correctly without the undisclosed operator, which is not included in this repository β€” loading it into a softmax model yields meaningless output; please do not benchmark that as "Tetracta rational."

The point of all this

This archive documents how a control overturned our initial interpretation. It is not current product validation. Negative controls can challenge an apparent effect; a single internal comparison cannot establish an architecture effect, and a recorded control does not make unrelated claims valid. The 12 September remeasurement above is a new, artifact-bound research execution with explicit controls and limitations. It does not reinstate the historical interpretations or establish ordinary customer support.

Tetracta AI Teams β€” for humans, like humans.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Space using tetracta/llm-xray-twin-study-1b 1

Collection including tetracta/llm-xray-twin-study-1b