DQN-Labs commited on
Commit
d0cc720
Β·
verified Β·
1 Parent(s): 7188484

Create README.md

Browse files
Files changed (1) hide show
  1. README.md +126 -0
README.md ADDED
@@ -0,0 +1,126 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language:
4
+ - en
5
+ base_model:
6
+ - Qwen/Qwen2.5-1.5B-Instruct
7
+ tags:
8
+ - dqnlabs
9
+ - coding
10
+ - python
11
+ - humaneval
12
+ - lora
13
+ metrics:
14
+ - name: HumanEval (zero-shot)
15
+ type: pass@1
16
+ value: 49.39%
17
+ ---
18
+
19
+ # 🧠 DQN Code v0.2
20
+ ### A 1.5B Python Specialist
21
+
22
+ DQN Code v0.2 is a lightweight coding-focused model built on **Qwen2.5-1.5B-Instruct** and fine-tuned specifically for high-quality Python generation.
23
+
24
+ This release focuses on **algorithmic correctness, structured implementation, and clean function completion**.
25
+
26
+ ---
27
+
28
+ ## πŸš€ Highlights
29
+
30
+ - πŸ”Ή 1.5B parameters
31
+ - πŸ”Ή LoRA fine-tuned
32
+ - πŸ”Ή Python-specialized
33
+ - πŸ”Ή Optimized for deterministic completion
34
+ - πŸ”Ή Designed to run locally on 8GB systems
35
+
36
+ ---
37
+
38
+ ## πŸ“Š Benchmark Performance
39
+
40
+ **HumanEval (0-shot, temperature=0.0)**
41
+ Evaluated using `mlx_lm`.
42
+
43
+ | Model | Parameters | HumanEval pass@1 |
44
+ |--------|------------|------------------|
45
+ | Qwen2.5-1.5B-Instruct (official) | 1.5B | 37.8% |
46
+ | Qwen2.5-Coder-1.5B (reported) | 1.5B | ~41.6% |
47
+ | Qwen2.5-Coder-1.5B (refined) | 1.5B | ~46.8% |
48
+ | **DQN Code v0.2** | 1.5B | **49.39%** |
49
+
50
+ > Evaluation settings:
51
+ > - 0-shot
52
+ > - Temperature = 0.0
53
+ > - No few-shot prompting
54
+ > - Full 164 HumanEval tasks
55
+
56
+ This represents a **+11.6% absolute improvement** over the base Instruct model.
57
+
58
+ ---
59
+
60
+ ## 🎯 Design Philosophy
61
+
62
+ Instead of scaling parameters, DQN Code focuses on:
63
+
64
+ - High-quality distilled supervision
65
+ - Python-heavy training distribution
66
+ - Clean function-style completions
67
+ - Reduced conversational overhead
68
+ - Local inference efficiency
69
+
70
+ Small models benefit heavily from specialization.
71
+ This release demonstrates how targeted fine-tuning can significantly improve coding performance without increasing model size.
72
+
73
+ ---
74
+
75
+ ## πŸ”§ Training Details
76
+
77
+ - Base: `Qwen2.5-1.5B-Instruct`
78
+ - Fine-tune type: LoRA
79
+ - Effective batch size: 8
80
+ - Max sequence length: 512
81
+ - Optimizer: AdamW
82
+ - Learning rate: 5e-6
83
+ - Dataset size: ~1.8k curated Python-focused samples
84
+ - Training hardware: 8GB RAM system (MLX)
85
+
86
+ Training focused on:
87
+ - Function completion
88
+ - Algorithmic correctness
89
+ - Clean Python structure
90
+ - Reduced hallucinated commentary
91
+
92
+ ---
93
+
94
+ ## πŸ’» Intended Use
95
+
96
+ - Local coding assistant (with decent performance in other languages too!)
97
+ - Python function completion
98
+ - Algorithm practice
99
+ - Educational use
100
+ - Lightweight code generation
101
+
102
+ ---
103
+
104
+ ## ⚠ Limitations
105
+
106
+ - Primarily optimized for Python (but performs well on other languages too.)
107
+ - Not benchmarked on multi-language coding
108
+ - Limited evaluation on mathematical reasoning
109
+ - Not trained for tool use or multi-step planning
110
+
111
+ ---
112
+
113
+ ## 🌍 Philosophy
114
+
115
+ Powerful coding models do not require massive infrastructure. You don't need a datacenter at home!
116
+
117
+ Focused training + efficient inference can deliver strong results on modest hardware.
118
+
119
+ ---
120
+
121
+ ## ❓ Queries
122
+ If you have any questions regarding the model, want to know how it was trained and our pipleine process, how you can make a better version of the model, or you just want to chat about AI, feel free to contact me on Discord at @dqnlabs.
123
+
124
+ Enjoy this model, for this is the *best* so far. There's more coming.
125
+
126
+ - **DQN Labs**