amine-yagoub commited on
Commit
c30312e
Β·
1 Parent(s): 8cf1c77

docs: enhance README with frontmatter and add Hugging Face version

Browse files
Files changed (2) hide show
  1. README.md +57 -16
  2. README_HF.md +293 -0
README.md CHANGED
@@ -1,5 +1,9 @@
 
1
  ## <<<<<<< HEAD
2
 
 
 
 
3
  title: CodeTribunal
4
  emoji: πŸ’»
5
  colorFrom: pink
@@ -8,18 +12,25 @@ sdk: docker
8
  pinned: false
9
  license: mit
10
  short_description: The AI Courtroom That Exposes Bad Freelance Code
 
11
 
12
  ---
13
 
14
  # Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
15
 
 
 
 
 
16
  <div align="center">
17
 
18
  # CodeTribunal
19
 
20
- ### The AI Courtroom That Exposes Bad Freelance Code
21
 
22
- **Built using GLM 5.1 for long-horizon reasoning and orchestrated tool use via CrewAI.**
 
 
23
 
24
  [![Tests](https://github.com/amineyagoub/CodeTribunal/actions/workflows/tests.yml/badge.svg)](https://github.com/amineyagoub/CodeTribunal/actions/workflows/tests.yml)
25
  [![Python 3.11+](https://img.shields.io/badge/python-3.11+-blue.svg)](https://www.python.org/downloads/)
@@ -30,18 +41,57 @@ short_description: The AI Courtroom That Exposes Bad Freelance Code
30
 
31
  ---
32
 
33
- ## The Problem
 
 
 
 
 
 
 
 
 
 
 
 
 
 
34
 
35
- A freelancer delivers code. The client can't tell if it's professional work or a security nightmare. Traditional linters find syntax errors. Code reviews miss architectural flaws. Nobody puts it all together and tells you:
36
 
37
- > _"This code is negligent, here's exactly why, and here's what it will cost you."_
38
 
39
- **CodeTribunal does.**
 
 
 
40
 
41
- Upload a `.zip` of code and watch a full courtroom trial unfold β€” evidence gathering, investigation by specialist agents, a live-streamed debate between an AI Prosecutor and Defense Attorney, and a Judge's verdict with a reputational risk score.
42
 
43
  ---
44
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
45
  ## How It Works
46
 
47
  CodeTribunal runs a **6-phase pipeline**, each building on the last:
@@ -280,15 +330,6 @@ Test fixtures in `tests/fixtures/locale/` contain deliberately bad Python and Ja
280
 
281
  ---
282
 
283
- ## What Makes This a Strong Hackathon Entry
284
-
285
- 1. **System Complexity** β€” 6-phase pipeline with 8 agents, 4 custom tools, code graph, and streaming
286
- 2. **Effective Tool Use** β€” Agents use BaseTool with Pydantic schemas to read files, search patterns, and trace calls
287
- 3. **Context Handoffs** β€” CrewAI `context` parameter chains prosecution β†’ defense β†’ rebuttal
288
- 4. **Custom ReACT Engine** β€” Direct LiteLLM function calling with GLM-5 for reliable tool use
289
- 5. **Deterministic + AI** β€” GritQL provides ground-truth evidence, agents provide interpretation and debate
290
- 6. **Resilient** β€” Rate-limit retry, pipeline persistence, error recovery
291
-
292
  ---
293
 
294
  <div align="center">
 
1
+ <<<<<<< HEAD
2
  ## <<<<<<< HEAD
3
 
4
+ =======
5
+ ---
6
+ >>>>>>> cc6493d (docs: enhance README with frontmatter and add Hugging Face version)
7
  title: CodeTribunal
8
  emoji: πŸ’»
9
  colorFrom: pink
 
12
  pinned: false
13
  license: mit
14
  short_description: The AI Courtroom That Exposes Bad Freelance Code
15
+ <<<<<<< HEAD
16
 
17
  ---
18
 
19
  # Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
20
 
21
+ =======
22
+ ---
23
+
24
+ >>>>>>> cc6493d (docs: enhance README with frontmatter and add Hugging Face version)
25
  <div align="center">
26
 
27
  # CodeTribunal
28
 
29
+ ### Put Freelance Code on Trial.
30
 
31
+ **Upload code. Get a verdict. Know the risk.**
32
+
33
+ Built with **GLM 5.1 + CrewAI + GritQL**
34
 
35
  [![Tests](https://github.com/amineyagoub/CodeTribunal/actions/workflows/tests.yml/badge.svg)](https://github.com/amineyagoub/CodeTribunal/actions/workflows/tests.yml)
36
  [![Python 3.11+](https://img.shields.io/badge/python-3.11+-blue.svg)](https://www.python.org/downloads/)
 
41
 
42
  ---
43
 
44
+ ## 🚨 The Problem
45
+
46
+ Clients receive code they don’t understand.
47
+
48
+ - Looks clean… but hides security risks
49
+ - Passes linters… but fails in production
50
+ - Works… but is architecturally broken
51
+
52
+ **No one answers the only question that matters:**
53
+
54
+ > _Is this code safe, professional, and worth paying for?_
55
+
56
+ ---
57
+
58
+ ## The Solution
59
 
60
+ **CodeTribunal turns code review into a courtroom trial.**
61
 
62
+ Upload a `.zip` β†’ get:
63
 
64
+ - Forensic evidence (AST-level)
65
+ - Multi-agent investigation
66
+ - AI courtroom debate
67
+ - Final verdict + risk score
68
 
69
+ > Not just analysis β€” **judgment**.
70
 
71
  ---
72
 
73
+ ## 🧠 Why This Exist
74
+
75
+ ### 1. Real System
76
+
77
+ - 6-phase pipeline
78
+ - 8 specialized agents
79
+ - Persistent execution engine
80
+
81
+ ### 2. Agents That Actually Act
82
+
83
+ - File reads, pattern search, call tracing
84
+ - Real tool usage via function calling (not fake reasoning)
85
+
86
+ ### 3. Deterministic + AI Hybrid
87
+
88
+ - **GritQL = ground truth**
89
+ - **Agents = interpretation + argument**
90
+
91
+ ### 4. End-to-End Story
92
+
93
+ From raw code β†’ evidence β†’ debate β†’ verdict β†’ report
94
+
95
  ## How It Works
96
 
97
  CodeTribunal runs a **6-phase pipeline**, each building on the last:
 
330
 
331
  ---
332
 
 
 
 
 
 
 
 
 
 
333
  ---
334
 
335
  <div align="center">
README_HF.md ADDED
@@ -0,0 +1,293 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ title: CodeTribunal
3
+ emoji: πŸ’»
4
+ colorFrom: pink
5
+ colorTo: red
6
+ sdk: docker
7
+ pinned: false
8
+ license: mit
9
+ short_description: The AI Courtroom That Exposes Bad Freelance Code
10
+ ---
11
+
12
+ <div align="center">
13
+
14
+ # CodeTribunal
15
+
16
+ ### Put Freelance Code on Trial.
17
+
18
+ **Upload code. Get a verdict. Know the risk.**
19
+
20
+ Built with **GLM 5.1 + CrewAI + GritQL**
21
+
22
+ [![Tests](https://github.com/amineyagoub/CodeTribunal/actions/workflows/tests.yml/badge.svg)](https://github.com/amineyagoub/CodeTribunal/actions/workflows/tests.yml)
23
+ [![Python 3.11+](https://img.shields.io/badge/python-3.11+-blue.svg)](https://www.python.org/downloads/)
24
+ [![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
25
+ [![Built for GLM 5.1 Hackathon](https://img.shields.io/badge/Built%20for-GLM%205.1-ff69b4)](https://build-with-glm-5-1-challenge.devpost.com)
26
+
27
+ </div>
28
+
29
+ ---
30
+
31
+ ## 🚨 The Problem
32
+
33
+ Clients receive code they don’t understand.
34
+
35
+ - Looks clean… but hides security risks
36
+ - Passes linters… but fails in production
37
+ - Works… but is architecturally broken
38
+
39
+ **No one answers the only question that matters:**
40
+
41
+ > _Is this code safe, professional, and worth paying for?_
42
+
43
+ ---
44
+
45
+ ## The Solution
46
+
47
+ **CodeTribunal turns code review into a courtroom trial.**
48
+
49
+ Upload a `.zip` β†’ get:
50
+
51
+ - Forensic evidence (AST-level)
52
+ - Multi-agent investigation
53
+ - AI courtroom debate
54
+ - Final verdict + risk score
55
+
56
+ > Not just analysis β€” **judgment**.
57
+
58
+ ---
59
+
60
+ ## 🧠 Why This Exist
61
+
62
+ ### 1. Real System
63
+
64
+ - 6-phase pipeline
65
+ - 8 specialized agents
66
+ - Persistent execution engine
67
+
68
+ ### 2. Agents That Actually Act
69
+
70
+ - File reads, pattern search, call tracing
71
+ - Real tool usage via function calling (not fake reasoning)
72
+
73
+ ### 3. Deterministic + AI Hybrid
74
+
75
+ - **GritQL = ground truth**
76
+ - **Agents = interpretation + argument**
77
+
78
+ ### 4. End-to-End Story
79
+
80
+ From raw code β†’ evidence β†’ debate β†’ verdict β†’ report
81
+
82
+ ## How It Works
83
+
84
+ CodeTribunal runs a **6-phase pipeline**, each building on the last:
85
+
86
+ ### Phase 1: Forensic Evidence (Deterministic β€” No LLM)
87
+
88
+ GritQL scans the entire codebase with **17 forensic patterns** across security and quality domains:
89
+
90
+ | Domain | Patterns | Examples |
91
+ | ----------- | -------- | ---------------------------------------------------------------------------------------- |
92
+ | πŸ”΄ Security | 13 | Hardcoded secrets, `eval()`, SQL injection, `pickle.load()`, `os.system()`, weak hashing |
93
+ | 🟑 Quality | 4 | `TODO`, `FIXME`, `HACK` comments |
94
+
95
+ All scanning is **read-only** (`--dry-run`) and runs in **parallel** across patterns.
96
+
97
+ ### Phase 2: Code Dependency Graph (AST β€” No LLM)
98
+
99
+ Python's `ast` module and regex-based JS parsing build a **lightweight dependency graph**:
100
+
101
+ - Nodes: files, functions, classes, imports
102
+ - Edges: calls, imports, containment, inheritance
103
+ - Enables call-chain tracing: `eval() β†’ handle_request() β†’ app.route()`
104
+
105
+ ### Phase 3: Investigation (3 ReACT Agents + 4 Tools)
106
+
107
+ Three specialist investigators, each running a **genuine ReACT loop** (Reason β†’ Act β†’ Observe β†’ Repeat) using **Z.ai's native function calling** via LiteLLM:
108
+
109
+ | Agent | Tools | Purpose |
110
+ | ------------------------- | --------------------------------------------------------- | ------------------------------------------ |
111
+ | Security Investigator | FileReader, PatternSearch, CodeGraphQuery, FindingContext | Find vulnerabilities, trace attack vectors |
112
+ | Quality Investigator | FileReader, FindingContext | Assess technical debt, detect negligence |
113
+ | Architecture Investigator | FileReader, CodeGraphQuery | Analyze structure, trace dependencies |
114
+
115
+ Each agent **autonomously decides which tools to call**, observes the results, and iterates. For example, the Security Investigator might:
116
+
117
+ 1. Call `file_reader` to read a flagged file
118
+ 2. Observe hardcoded secrets on specific lines
119
+ 3. Call `code_graph_query` to trace where those secrets are used
120
+ 4. Produce a detailed report with file paths, line numbers, and severity ratings
121
+
122
+ **Verified working**: GLM-5 + LiteLLM function calling confirmed. Agents make real tool calls that execute real code analysis.
123
+
124
+ ### Phase 4: The Trial (3 Agents)
125
+
126
+ A courtroom debate between AI agents:
127
+
128
+ 1. ** The Prosecutor** β€” builds the case for negligence, cites specific evidence
129
+ 2. ** The Defense Attorney** β€” challenges claims, argues context and proportionality
130
+ 3. ** Rebuttal** β€” the prosecutor responds to the defense
131
+
132
+ Agents use CrewAI's `context` parameter to chain arguments: prosecution output feeds into defense context, both feed into rebuttal.
133
+
134
+ ### Phase 5: The Verdict
135
+
136
+ ** The Judge** reviews all evidence, investigation reports, and the full trial transcript. Delivers:
137
+
138
+ - Overall ruling: GUILTY / MIXED / NOT GUILTY
139
+ - Reputational Risk Score (0-100)
140
+ - Findings summary with severity rankings
141
+
142
+ ### Phase 6: Structured Report
143
+
144
+ ** Verdict Report Agent** compiles everything into a professional report:
145
+
146
+ - Executive Summary
147
+ - Findings Table (sorted by severity)
148
+ - Per-Finding Analysis (impact, remediation, estimated fix effort)
149
+ - Sentencing Recommendations
150
+
151
+ ---
152
+
153
+ ## Architecture
154
+
155
+ ```
156
+ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
157
+ β”‚ Gradio UI β”‚
158
+ β”‚ + Export β”‚
159
+ β””β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”˜
160
+ β”‚
161
+ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
162
+ β”‚ Pipeline Engine β”‚
163
+ β”‚ State Β· Persistence β”‚
164
+ β”‚ Cancel Β· Resume β”‚
165
+ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
166
+ β”‚
167
+ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
168
+ β–Ό β–Ό β–Ό β–Ό β–Ό
169
+ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”
170
+ β”‚Evidence β”‚ β”‚Code β”‚ β”‚Invest. β”‚ β”‚ Trial β”‚ β”‚Reportβ”‚
171
+ β”‚ Scanner β”‚ β”‚Graph β”‚ β”‚ Agents β”‚ β”‚ Agents β”‚ β”‚Agent β”‚
172
+ β”‚(GritQL) β”‚ β”‚(AST) β”‚ β”‚+ Tools β”‚ β”‚ β”‚ β”‚ β”‚
173
+ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”˜
174
+ β”‚ β”‚ β”‚ β”‚ β”‚
175
+ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
176
+ β”‚
177
+ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
178
+ β”‚ Custom Tool Layer β”‚
179
+ β”‚ FileReader Β· Pattern β”‚
180
+ β”‚ CodeGraph Β· Context β”‚
181
+ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
182
+ ```
183
+
184
+ ### Key Design Decisions
185
+
186
+ | Decision | Why |
187
+ | ------------------------------------- | ------------------------------------------------------------------------------------------- |
188
+ | **Agents have tools, not text dumps** | Agents read files, search patterns, and trace calls on demand β€” scales to any codebase size |
189
+ | **ReACT loop via LiteLLM** | Direct function calling with GLM-5 β€” bypasses CrewAI's unreliable tool routing |
190
+ | **Pipeline state persisted to JSON** | Runs can resume after crashes. State is queryable |
191
+ | **GritQL for evidence** | AST-level pattern matching, not regex. Language-aware, precise |
192
+ | **Custom CrewAI tools (BaseTool)** | Pydantic-validated inputs, proper error handling, CrewAI-native integration |
193
+ | **Rate-limit retry with backoff** | Exponential backoff (4s β†’ 64s) on Z.ai 429 errors β€” pipeline survives API spikes |
194
+
195
+ ---
196
+
197
+ ## Tech Stack
198
+
199
+ | Component | Technology | Purpose |
200
+ | -------------------- | -------------------------- | ---------------------------------------------------- |
201
+ | **LLM** | GLM 5.1 via Z.ai (LiteLLM) | Agent reasoning and debate |
202
+ | **Code Scanning** | GritQL | Deterministic AST-level pattern matching |
203
+ | **Multi-Agent** | CrewAI 1.12 | Agent orchestration, task chaining, context handoffs |
204
+ | **Function Calling** | LiteLLM | Direct ReACT loop with GLM-5 tool calling |
205
+ | **Code Graph** | Python `ast` + regex | Dependency graph (Python + JS) |
206
+ | **UI** | Gradio 6 | Streaming chatbot, file upload, export |
207
+ | **Export** | markdown-pdf (PyMuPDF) | PDF report generation from Markdown |
208
+
209
+ ---
210
+
211
+ ## Install
212
+
213
+ ```bash
214
+ # Clone
215
+ git clone https://github.com/amineyagoub/CodeTribunal.git
216
+ cd CodeTribunal
217
+
218
+ # Install dependencies
219
+ pip install -e .
220
+
221
+ # Install GritQL CLI
222
+ npm install -g @getgrit/cli
223
+
224
+ # Configure
225
+ cp .env.example .env
226
+ # Edit .env: set ZAI_API_KEY (get one at https://open.bigmodel.cn/)
227
+ ```
228
+
229
+ ### Requirements
230
+
231
+ - Python 3.11+
232
+ - Node.js (for GritQL CLI)
233
+ - Z.ai API key ([get one here](https://open.bigmodel.cn/))
234
+
235
+ ---
236
+
237
+ ## Usage
238
+
239
+ ### Web UI (Recommended)
240
+
241
+ ```bash
242
+ python3 -m code_tribunal.app
243
+ ```
244
+
245
+ Open http://localhost:7860, upload a `.zip` of code, and watch the trial unfold.
246
+
247
+ ---
248
+
249
+ ## Features
250
+
251
+ | Feature | Details |
252
+ | ------------------------------ | --------------------------------------------------------------------------------------- |
253
+ | **4 Custom Tools** | FileReader, PatternSearch, CodeGraphQuery, FindingContext β€” agents actively investigate |
254
+ | **8 Specialized Agents** | 3 investigators, prosecutor, defense, rebuttal, judge, verdict report, expert witness |
255
+ | **ReACT Engine** | Custom Reason-Act-Observe loop via LiteLLM function calling with GLM-5 |
256
+ | **Code Dependency Graph** | AST-based (Python + JS), with call-chain tracing and impact analysis |
257
+ | **Parallel Evidence Scanning** | ThreadPoolExecutor for GritQL patterns β€” 4x faster than sequential |
258
+ | **Rate-Limit Resilience** | Exponential backoff retry on 429 errors β€” survives API rate limits |
259
+ | **Pipeline Persistence** | State saved to JSON, runs can resume after interruption |
260
+ | **Deduplication** | Same file+line merged into one finding with multiple categories |
261
+ | **Zip Safety** | Zip-slip attack prevention |
262
+ | **Streaming UI** | Real-time pipeline progress in Gradio Chatbot with phase indicators |
263
+ | **Export** | Markdown and PDF report generation |
264
+
265
+ ---
266
+
267
+ ## πŸ§ͺ Testing
268
+
269
+ ```bash
270
+ # Run evidence scan on test fixtures
271
+ code-tribunal tests/fixtures/locale/ --evidence-only
272
+
273
+ # Run Python tests
274
+ pytest tests/
275
+ ```
276
+
277
+ Test fixtures in `tests/fixtures/locale/` contain deliberately bad Python and JavaScript code with:
278
+
279
+ - Hardcoded passwords, API keys, AWS secrets, Stripe keys, JWT secrets
280
+ - SQL injection via f-strings and template literals
281
+ - `eval()`, `pickle.load()`, `os.system()`, `subprocess.call(shell=True)`
282
+ - MD5 hashing
283
+ - TODO, FIXME, HACK comments
284
+
285
+ ---
286
+
287
+ ---
288
+
289
+ <div align="center">
290
+
291
+ Built for the [Build with GLM 5.1](https://build-with-glm-5-1-challenge.devpost.com) hackathon.
292
+
293
+ </div>