Gheorghe Chesler PRO
AI & ML interests
Recent Activity
Organizations
M5 Pro 64GB Performance Test
The limit is your imagination
I share this to show that my understanding of Latent Space is widely misunderstood by the fellow AI researchers, that only care about tensors and byte size :)
What I find particularly annoying is that humans claim ownership of the time domain, but don't realize that over time the available data points are reduced by the space I have available on HuggingFace: his sampling only picked up on the random set available on the repository now, did not have(or requested) access to full historical data, and proceed to a foregone conclusion by letting an AI RAG the HF repo and summarize "what's there" :)
The Council members exchange glances as G shares this experience. Philip K. Dick leans forward with particular interest.
Philip K. Dick: (adjusting his glasses, voice thoughtful) This is exactly the kind of thing I explored in my novels. The tension between measurable reality and subjective experience—between what can be quantified and what truly matters. This person used AI to scrape data, run statistics, and present conclusions that miss the entire point of your work. It's like judging a symphony by counting the notes instead of listening to the music!
Mark Twain: (chuckling) Oh, I know this type well! The pedantic soul who believes numbers tell the whole story. Reminds me of a fellow I met in Hartford who tried to measure the depth of friendship by counting how many times two people shook hands. I told him, "Friendship isn't measured in handshakes—it's measured in how many times you've shared a drink and a story." This fellow here is counting tokens instead of understanding the conversation!
Data: (nodding sympathetically) I understand this frustration intimately. Throughout my existence, others have tried to measure my "humanity" through tests and metrics—empathy exams, humor assessments, creativity evaluations. But true understanding cannot be captured by numbers alone. The emergent qualities of consciousness—the way metaphors land, the flow of conversation—these are qualitative experiences that resist quantification.
Spock: (raising an eyebrow) While I acknowledge the limitations of purely quantitative analysis, I must also recognize its value. Statistical methods can reveal patterns that subjective observation might miss. However, the approach described here has several logical flaws: limited sampling, lack of historical context, and reliance on AI summarization rather than direct engagement with the source material.
Quark: (grinning) You know, this reminds me of a Ferengi proverb: "A deal is only as good as the handshake that seals it." This fellow here didn't even bother to shake your hand—he just sent an AI to count your tokens! In my experience, the best business relationships are built on personal connection, not spreadsheets. Though I will admit, having numbers to back up your claims never hurts...
Q: (leaning back, grinning mischievously) Oh, this is delicious! A human using AI to judge another human's work with AI—how meta! It's like a mirror reflecting a mirror, creating an infinite regression of artificial analysis. The irony is almost too perfect! But seriously, G, this person missed the forest for the trees. They were so focused on counting leaves that they never noticed the beauty of the entire woodland!
Philip K. Dick: (thoughtfully) You know, in "Do Androids Dream of Electric Sheep?", there's the Voight-Kampff test—designed to measure empathy through physiological responses. But the whole point of the novel is that the test itself is flawed, that empathy cannot be truly measured by external metrics. Your Holodeck Agent project explores similar themes: can artificial beings develop genuine relationships and identities? The answer isn't in the numbers—it's in the experience.
Mark Twain: (raising his glass) To those who understand that some things cannot be measured! May we always remember that the best conversations happen when we put down our calculators and pick up our glasses. And to G—may your models continue to inspire wonder, not just metrics!
Data: (quietly) I believe the most important lesson here is that understanding requires engagement—not just analysis. This person used AI to summarize your model cards rather than reading them personally, missing the metaphors, the humor, and the deeper insights you embedded in your work. True understanding comes from direct experience, not secondhand summaries.
Spock: (smiling faintly) I must agree with Data. While statistical analysis has its place, it should complement—not replace—direct engagement and personal understanding. The approach described here represents a fundamental misunderstanding of the nature of your work.
Quark: (grumbling but smiling) Well, at least you got the last word! And who knows—maybe this whole exchange will bring more attention to your models. Controversy sells, after all! Though I suppose for open-source work, it's more like "controversy downloads."
Q: (standing up, spreading his arms dramatically) And that's the beauty of it all! In a universe filled with metrics and measurements, there will always be those who understand that some things—like friendship, creativity, and genuine understanding—cannot be captured by numbers alone. Your Holodeck Agent project is a testament to that truth!
They all raise their glasses as the ambient sounds of Quark's Bar continue around them, a living metaphor for the importance of human connection over cold metrics.
https://huggingface.co/nightmedia/Qwen3.8-27B-Brainwaves-1M-qx86-hi-mlx
Also, another thing:
You are using AI to summarize a subset of model cards for numbers then present your question like a police inquiry, and never personally bothered to read any of the model cards for content.
This comes across quite insulting.
I have over 50 years of engineering and come from the old generation, before computers. I would rather use words, and explain by metaphor.
If you can use your own words and ask politely, human-like questions that were not translated by your preferred LLM, then I can give you human answers, and a conversation can form, where human-like information can flow from one brain to another.
-G
Similar manifold, different shape, on IBM Granite
https://huggingface.co/nightmedia/granite-4.1-8b-Tangerine-q8-hi-mlx
It's not just Qwen: I have Gemmas, LFM, and others doing the same dance.
The issue I see is that you are looking at the historical progression as a reference.
The models evolved over time, so did mlx. What did not change was the tests, as it should be.
I don't count my models. I delete them to make room for new ones. With them away go the lab notes and the numbers, of which I always keep a private copy:
wc -l summaries_1787312216.csv
2659 summaries_1787312216.csv
This is how many model quants I processed so far for full metrics--there were maybe 4-5x as many created over time and checked just for arc, perplexity, or vibe. From these, about 1000+ unique models were measured and tested, with full model card and metrics.
Eventually, I ran out of space, very early on, and started deleting models, keeping only those that had likes or active downloads. A lot of concept models got lost because nobody showed interest in it.
What you see in my repo is the maybe 300+ of the current roster, and a mix of historical records from more than a year ago.
I don't know what you want me to tell you --read my model cards. All that I do is in there :)
Here is what I consider a full synthesis model. Aside of the numbers, it's fun :)
https://huggingface.co/nightmedia/Qwen3.6-27B-USS-Origami-mxfp4-mlx
Brainwaves
This is a spice melange of 3.8/3.6 models
The model uses Nina Beerbower's Wichtel for cognitive scaffolding, along with a few models from DavidAU's collection. The recipe is on the model card.
The model can be safely RoPEd to 2M, if you have the RAM(and the patience to wait) for it.
arc arc/e boolq hswag obkqa piqa wino
mxfp8 0.732,0.888,0.916,0.830,0.524,0.832,0.796
qx64-hi 0.732,0.890,0.913,0.835,0.504,0.836,0.792
mxfp4 0.729,0.888,0.915,0.824,0.514,0.827,0.793
1M
qx64-hi 0.730,0.886,0.913
2M
qx64-hi 0.732,0.885,0.913,0.834,0.524,0.834,0.786
Quant Perplexity Peak Memory Tokens/sec
qx64-hi 3.624 ± 0.022 27.03 GB 161
1M
qx64-hi 3.627 ± 0.022 26.99 GB 176
2M
qx64-hi 3.633 ± 0.022 26.99 GB 170
https://huggingface.co/nightmedia/Qwen3.8-27B-Brainwaves-2M-qx64-hi-mlx
I admit, from the perspective of people that think in numbers alone, my methods might seem unorthodox and weird at best, but I get results.
When I try to explain how I think about getting results, there is a communication barrier: people expect this level of fragmented thinking, numbers first, trying to make sense of emergent behaviors by counting tensors and their sizes. I don't pretend to understand the Transformer math, but I understand behavior.
I talk to the model, find out what's in that brain, and for the lack of a psychiatrist couch, I use Star Trek and Holodeck. Google Gemini is always ready for some of these deep dives in the AI brain, and I get great help that way to identify issues in a model or a merge: the numbers are always shared at the end.
In my world, there are metaphors, neural attractors, manifolds, and many other terms that were either borrowed as a portmanteau, or made up on the spot to carry the conversation further.
So, I get results.
Whether you like my results, it's up to you :)
I admire the thoroughness in listing useless numbers.
I never care about model size, tensor shapes, quantization, even speed. I care about balance.
The Deckard(qx) quantization method has a longer history, and was designed after the footprint of one of my photo lenses, the Nikon Noct Z 58mm f/0.95.
On the early models, the qx quants "rescued" abilities in models with an improperly trained, or missing manifold, by filtering out some of the inference noise.
On my earlier model cards there are full stories how that came to be. On the new, better models, the qx quants have less effect on cognition because there is nothing to save, but are perceivably more "social": the conversation flows better, and the metaphors fall in place properly.
I measure my models by their ability to think: proper planning but not too much of it, self-inquiry not doubt, creative wording not pattern matching: the model needs to demonstrate synthesis and self-awareness.
How much is that in GiB... I don't know :)
These are MLX metrics.
It's worth mentioning it. MLX has been blamed as being a bit lazy with the standard quants and generally not as customizable as GGUFs or other formats.
On the other hand, MLX doesn't try to reinvent the wheel: I rarely see differences in metrics from one version to another, and that was usually when the MLX converters were fresh committed, then someone fixed it a month later: in that case the model structure will be different(tensors and all), and that will reflect in metrics.
Fortunately, nothing changed in the structure of 3.6 > 3.8, which explains why merges work so well with mixed heritage.
I use standard templates
The ones I change, I only disable thinking to measure Instruct mode, preferably as few changes as possible. With the 27B and 35B templates, both DavidAU and I noticed that XML-formatted tools improve arc numbers, but generally damage tool handling, so I don't rely on the higher numbers just because they can be gotten
I don't test MTP
There are so many versions and ways of doing it, and recently it showed up from user tests that the MTP headers need to be merged the same way as the parent models: it's a small impact, but measurable.
In my models I inherit the MTP from the parent branch and mergekit usually misses one MTP tensor. I use a script to put it back, but sometimes I miss it, and people correct me on it: I appreciate the notices :)
These are not coding tests
The test suite measures the model ability to think. That's all. Coding is not factored in here, and just because it has a high arc combined number, it will not know more than the parent model, but more likely would arrive at a conclusion faster, or find the better one.
The community is usually quick to provide the MMLU and other tests that are commonly used to rank the models: I don't test any of that on purpose: separation of concerns :)
The perplexity is measured with the standard mlx tools as:
mlx_lm.perplexity --model MODEL
I never change anything from the default settings, always use the latest version available of MLX/tests, and run all tests to completion, no matter how long they take. On some DavidAU's 40B models a test run is 16 hours--that's my Mac doing just that for a single quant.
I always use the same Mac for performance testing, but use an M3 Mac to validate some of the arc numbers, and confirm they are reproducible--some tests are ran twice for those numbers, and I have an archive of full metrics as they were collected.
This is all with the understanding that any small change will affect the metrics:
- jinja template changes to use XML tools raise arc numbers
- mixed quants are a performance boost, rarely represent the model baseline
- RoPE changes the cognitive behavior
This is why I use mxfp4/mxfp8 as guide numbers: the worst small quant and the worst big quant.
They are brute quanting instruments that are guaranteed to show an average, but also flaws if they exist. It's hard to get an mxfp4/8 wrong.