AhiskaAI-65m-IT-v0.2

AhiskaAI-65m-IT-v0.2 is the instruction-tuned version of our 65M parameter Small Language Model. Fine-tuned on a curated Turkish instruction dataset, it is designed to function as a lightweight conversational AI assistant while maintaining fast inference on resource-constrained hardware.

Base Model: AhiskaAI-65m-Base-v0.2


Model Details

  • Architecture: Llama-based architecture.
  • Fine-tuning: Supervised Fine-Tuning (SFT).
  • Format: ChatML.
  • Parameters: 65M.
  • Context Window: 1024 tokens.
  • Tokenizer: Custom BPE Tokenizer (Vocabulary Size: 32,000).
  • Training Framework: PyTorch & Transformers.
  • Hardware: NVIDIA RTX 4050 6GB Laptop GPU.

Fine-tuning Dataset

The model was fine-tuned using a curated Turkish instruction dataset designed to improve conversational ability and instruction following.

The dataset focuses on:

  • Question answering
  • General conversation
  • Summarization
  • Text generation
  • Turkish instruction following

Design Goal

The 65M-IT model is designed as the lightweight conversational member of the AhiskaAI v0.2 family.

Its primary goals are:

  • Basic Turkish instruction following.
  • Fast conversational inference.
  • Low-resource deployment.
  • A compact research baseline for future alignment methods.

Training Logs

Training Loss Curve

The graph above demonstrates the supervised fine-tuning convergence of AhiskaAI-65m-IT-v0.2.


Usage (ChatML)

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("AhiskaAI/AhiskaAI-65m-IT-v0.2")
tokenizer = AutoTokenizer.from_pretrained("AhiskaAI/AhiskaAI-65m-IT-v0.2")

SYSTEM_PROMPT = "Sen kibar, sorulan soruları tam cümlelerle yanıtlayan Türkçe bir asistansın."

prompt = (
    f"<|im_start|>system\n{SYSTEM_PROMPT}<|im_end|>\n"
    f"<|im_start|>user\nMerhaba<|im_end|>\n"
    f"<|im_start|>assistant\n"
)

inputs = tokenizer(prompt, return_tensors="pt")

outputs = model.generate(**inputs, max_new_tokens=200)

print(tokenizer.decode(outputs[0], skip_special_tokens=True))

Known Limitations

  • Limited factual knowledge due to model size.
  • Optimized primarily for Turkish.
  • Context window limited to 1024 tokens.
  • May generate inaccurate or incomplete responses on complex topics.

Future Plans

  • DPO preference alignment.
  • Improved instruction datasets.
  • Future AhiskaAI v0.3 releases.

About AhiskaAI

AhiskaAI is an independent open-source initiative dedicated to developing efficient Turkish Small Language Models trained completely from scratch.

Follow us on Hugging Face for updates and future releases.

Downloads last month
40
Safetensors
Model size
70.7M params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including AhiskaAI/AhiskaAI-65m-Instruct-v0.2