How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf mervinpraison/Phi-4-harupfall-axis:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf mervinpraison/Phi-4-harupfall-axis:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf mervinpraison/Phi-4-harupfall-axis:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf mervinpraison/Phi-4-harupfall-axis:Q4_K_M
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf mervinpraison/Phi-4-harupfall-axis:Q4_K_M
# Run inference directly in the terminal:
./llama-cli -hf mervinpraison/Phi-4-harupfall-axis:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf mervinpraison/Phi-4-harupfall-axis:Q4_K_M
# Run inference directly in the terminal:
./build/bin/llama-cli -hf mervinpraison/Phi-4-harupfall-axis:Q4_K_M
Use Docker
docker model run hf.co/mervinpraison/Phi-4-harupfall-axis:Q4_K_M
Quick Links

Uploaded model

  • Developed by: mervinpraison
  • Finetuned from model : unsloth/Phi-4-unsloth-bnb-4bit
dataset:
- name: mervinpraison/harup-fall-axis-alpaca
dataset_num_proc: 2
dataset_text_field: text
gradient_accumulation_steps: 2
hf_model_name: mervinpraison/Phi-4-harupfall-axis
huggingface_save: 'true'
learning_rate: 0.0001
load_in_4bit: true
loftq_config: null
logging_steps: 15
lora_alpha: 16
lora_bias: none
lora_dropout: 0
lora_r: 16
lora_target_modules:
- q_proj
- k_proj
- v_proj
- o_proj
- gate_proj
- up_proj
- down_proj
lr_scheduler_type: linear
max_seq_length: 2048
max_steps: 6000
model_name: unsloth/Phi-4-unsloth-bnb-4bit
model_parameters: 14b
num_train_epochs: 10
ollama_model: mervinpraison/Phi-4-harupfall-axis
ollama_save: 'true'
optim: lion_8bit
output_dir: outputs
packing: false
per_device_train_batch_size: 1
quantization_method:
- q4_k_m
random_state: 3407
seed: 3407
train: 'true'
use_gradient_checkpointing: unsloth
use_rslora: false
warmup_steps: 100
weight_decay: 0.05

Training Details: wandb

Downloads last month
9
Safetensors
Model size
15B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support