--- license: mit tags: - llama.cpp - cuda - colab - t4 - sm75 --- # llama.cpp b10355 CUDA runtime for Colab T4 (SM75) This public repository stores a reproducible `llama-server` runtime, not model weights. It was cross-built on an actual Google Colab G4 with CUDA 12.8 and `CMAKE_CUDA_ARCHITECTURES=75`, then downloaded without authentication and validated on an actual Colab Tesla T4. ## Provenance - Source: `ggml-org/llama.cpp` - Release: `b10355` - Commit: `dd1ea524333b1e697489067d7a4c39c60d32beee` - Build host: NVIDIA RTX PRO 6000 Blackwell Server Edition (Colab G4) - Target: Linux x86-64, CUDA 12.8, SM75, Colab Tesla T4 - Archive: `llama-cpp-b10355-cuda128-sm75-colab.tar.gz` - Archive SHA-256: `4469c6f18805f2cf2917df6bb8c0dbd37937344ab6c2e77c312cccdf7b145a9a` - `llama-server` SHA-256: `f4f421a690bc949174752111bd20ac861fcdc2ba65fca361955edf7c67b53252` ## Verified T4 result The public archive was restored at revision `5738166d5979c9d39d0b3cacbaa1c5c2a3d44b16` with no HF token. It loaded `unsloth/Qwen2.5-VL-7B-Instruct-GGUF` revision `68bb8bc4b7df5289c143aaec0ab477a7d4051aab`, using `UD-Q4_K_XL` plus `mmproj-BF16`. - GPU memory after load: 6,647 MiB - Model load to healthy server: 4.01 seconds on the verified rerun - Fixed workload: uncached 512 prompt tokens to fixed 128 output tokens, batch 1, concurrency 1, five repetitions - Client end-to-end median: 34.53 tok/s - llama.cpp decode median: 39.20 tok/s - Five output hashes were identical - Multimodal smoke: 1024x768 candy photo, 4.00 seconds; correctly reported five candies and the teal, orange, and green colors See `t4_validation.json` and `t4_setup_receipt.json` for raw evidence. ## Restore ```python from huggingface_hub import hf_hub_download import hashlib, tarfile path = hf_hub_download( repo_id="WJO/llama-cpp-b10355-colab-t4-sm75", filename="llama-cpp-b10355-cuda128-sm75-colab.tar.gz", revision="5738166d5979c9d39d0b3cacbaa1c5c2a3d44b16", ) assert hashlib.sha256(open(path, "rb").read()).hexdigest() == ( "4469c6f18805f2cf2917df6bb8c0dbd37937344ab6c2e77c312cccdf7b145a9a" ) with tarfile.open(path, "r:gz") as archive: archive.extractall("/content", filter="data") ``` The runtime expects the CUDA and standard system libraries included in the current Colab CUDA 12.8 image. Verify the archive and binary hashes before use.