Instructions to use unsloth/GLM-4.7-Flash-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use unsloth/GLM-4.7-Flash-GGUF with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="unsloth/GLM-4.7-Flash-GGUF") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("unsloth/GLM-4.7-Flash-GGUF", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use unsloth/GLM-4.7-Flash-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf unsloth/GLM-4.7-Flash-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: llama cli -hf unsloth/GLM-4.7-Flash-GGUF:UD-Q4_K_XL
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf unsloth/GLM-4.7-Flash-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: llama cli -hf unsloth/GLM-4.7-Flash-GGUF:UD-Q4_K_XL
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf unsloth/GLM-4.7-Flash-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: ./llama-cli -hf unsloth/GLM-4.7-Flash-GGUF:UD-Q4_K_XL
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf unsloth/GLM-4.7-Flash-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: ./build/bin/llama-cli -hf unsloth/GLM-4.7-Flash-GGUF:UD-Q4_K_XL
Use Docker
docker model run hf.co/unsloth/GLM-4.7-Flash-GGUF:UD-Q4_K_XL
- LM Studio
- Jan
- vLLM
How to use unsloth/GLM-4.7-Flash-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "unsloth/GLM-4.7-Flash-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "unsloth/GLM-4.7-Flash-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/unsloth/GLM-4.7-Flash-GGUF:UD-Q4_K_XL
- SGLang
How to use unsloth/GLM-4.7-Flash-GGUF with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "unsloth/GLM-4.7-Flash-GGUF" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "unsloth/GLM-4.7-Flash-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "unsloth/GLM-4.7-Flash-GGUF" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "unsloth/GLM-4.7-Flash-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use unsloth/GLM-4.7-Flash-GGUF with Ollama:
ollama run hf.co/unsloth/GLM-4.7-Flash-GGUF:UD-Q4_K_XL
- Unsloth Desktop
- Pi
How to use unsloth/GLM-4.7-Flash-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf unsloth/GLM-4.7-Flash-GGUF:UD-Q4_K_XL
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "unsloth/GLM-4.7-Flash-GGUF:UD-Q4_K_XL" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use unsloth/GLM-4.7-Flash-GGUF with Docker Model Runner:
docker model run hf.co/unsloth/GLM-4.7-Flash-GGUF:UD-Q4_K_XL
- Lemonade
How to use unsloth/GLM-4.7-Flash-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull unsloth/GLM-4.7-Flash-GGUF:UD-Q4_K_XL
Run and chat with the model
lemonade run user.GLM-4.7-Flash-GGUF-UD-Q4_K_XL
List all available models
lemonade list
- Hermes Agent
How to use unsloth/GLM-4.7-Flash-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf unsloth/GLM-4.7-Flash-GGUF:UD-Q4_K_XL
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default unsloth/GLM-4.7-Flash-GGUF:UD-Q4_K_XL
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use unsloth/GLM-4.7-Flash-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf unsloth/GLM-4.7-Flash-GGUF:UD-Q4_K_XL
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "unsloth/GLM-4.7-Flash-GGUF:UD-Q4_K_XL" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Getting 110 tokens/sec on my RTX 3090, 24 GB VRam
Build llama.cpp on WSL and running the GLM-4.7-Flash-UD-Q4_K_XL.gguf model.
Context size is 120K which can be comfortably fit on my vram. This gave me 110 tokens/sec.
I also tried GLM-4.7-Flash-UD-Q5_K_XL.gguf model which performed similarly at 60k context. But as the context fills the memory starts to gets occupied resulting in gradual speed decline (2-3 token / sec decline per turn)
Here is a rough benchmark between the 2 models.
Benchmark Question
Context
You are acting as a senior applied AI engineer asked to help a product team diagnose churn and data quality issues in a subscription service.
Below is a messy event log extract and a short product requirements note. The data is intentionally imperfect.
Product Notes (read carefully)
Users can have multiple subscriptions, but only one can be active at a time.
cancelled_atmay appear beforesubscription_startdue to backfilled data.Timestamps are in mixed time zones, but all end with
Zor an offset.The business defines 30‑day retention as:
A user who still has an active subscription 30 days after their first subscription start.
The analytics team wants:
A retention rate
A list of data quality issues
A clean, reusable function that could be productionized
Assume no external internet access, but standard Python libraries are allowed.
Event Log (CSV)
user_id,subscription_id,event_type,event_time u1,s1,started,2023-01-01T10:00:00Z u1,s1,cancelled,2023-01-20T09:00:00Z u2,s2,started,2023-01-05T08:00:00-05:00 u2,s3,started,2023-02-10T12:00:00Z u3,s4,started,2023-01-15T14:00:00Z u3,s4,cancelled,2023-01-10T10:00:00Z u4,s5,started,2023-01-01T00:00:00Z u4,s5,cancelled,2023-02-15T00:00:00Z u5,s6,started,2023-01-31T23:59:59Z
Your Tasks
1. Clarification & Assumptions
List up to 3 clarifying questions you would ask the product or data team.
Then proceed anyway, clearly stating the assumptions you choose if those questions remain unanswered.
2. Data Reasoning
Identify at least 5 non‑trivial data quality or modeling issues in the dataset.
Explain why each issue matters for retention analysis.
3. Coding
Write a clean, readable Python function that:
Parses the data
Normalizes timestamps
Computes the 30‑day retention rate
Include brief inline comments and describe the time complexity.
4. Validation & Testing
Propose 2–3 test cases (edge cases preferred).
Explain how you would validate correctness in production.
5. Agentic / Tool Reasoning
If you had access to tools (e.g., Python execution, SQL, dashboards), describe:
Which tools you would use
In what order
What signals would tell you to change course
6. Executive Summary
Write a 5–7 sentence summary for a non‑technical stakeholder explaining:
The retention result
Confidence level
Key risks and next steps
Constraints
Do not skip steps.
Do not answer with bullet points only—mix narrative, structure, and code.
Optimize for clarity, correctness, and judgment, not brevity.
| Metric | Response 1 Score GLM-4.7-Flash-UD-Q5_K_XL.gguf | Response 1 remarks | GLM-4.7-Flash-UD-Q4_K_XL.gguf | Response 2 remarks |
|---|---|---|---|---|
| Coding ability | 30 | Major failure: provided function crashes (KeyError: 'active') because it never records per-user active status but later sums cohorts['active']. Also misstates time complexity (loop filters DF per user ⇒ closer to O(U·N)). Uses pandas despite “standard libraries” constraint ambiguity. |
70 | Code is runnable and uses stdlib (csv, datetime), good portability. Timestamp parsing via fromisoformat (+ Z fix) is solid. Weaknesses: claims UTC normalization but doesn’t explicitly convert; ignores subscription_id and does not model “restart within 30 days” correctly; complexity statement mentions sorting but code doesn’t sort. |
| Multi-turn reasoning (agentic & tool use) | 78 | Strong: good clarifying questions, explicit assumptions, decent tool stack (SQL, Great Expectations, orchestration, dashboards) + signals. | 72 | Good structure and sensible tool ordering + signals. Slightly less concrete than R1; some assumptions are stated but then inconsistently applied later. |
| Raw intelligence | 70 | Catches several real issues (timezones, backfilled cancels, user vs sub aggregation). Some “issues” are more about parsing/library choice than modeling. Minor date reasoning glitch (mentions Feb 29 in 2023) but overall conclusion aligns with the intended logic. | 60 | Good identification of common pitfalls, but makes a clear retention reasoning error (calls u5 churned despite no cancellation). Some points are fluff (DST mention without relevance here). |
| Coherence over long context | 55 | Narrative is broadly consistent, but coherence breaks because the executive summary depends on a metric the code cannot compute (code crashes). Also vacillates about stdlib vs pandas. | 45 | Biggest issue: internal contradiction—code returns 80%, but the response repeatedly asserts 60% (comments + executive summary). Also says “sort events” while not actually sorting. |
| Production readiness & robustness (new) | 50 | Good instincts (monitoring, validation tools, pipeline audit), but not production-ready due to non-running code + dependency mismatch + weak outputs for “data quality issues” (just a string). | 65 | More production-friendly: stdlib-only, clear function signature, test ideas + SQL baseline validation plan. Still lacks subscription-state modeling and returns only a float (no issue list / audit output), but closer to shippable. |
| Overall (average of above) | 57 | Strong analysis/agent framing, but coding execution quality is a blocker. | 62 | Better executable code + portability, but major self-contradiction on the key result hurts trust. |