Spaces:
Running
Running
Commit Β·
c30312e
1
Parent(s): 8cf1c77
docs: enhance README with frontmatter and add Hugging Face version
Browse files- README.md +57 -16
- README_HF.md +293 -0
README.md
CHANGED
|
@@ -1,5 +1,9 @@
|
|
|
|
|
| 1 |
## <<<<<<< HEAD
|
| 2 |
|
|
|
|
|
|
|
|
|
|
| 3 |
title: CodeTribunal
|
| 4 |
emoji: π»
|
| 5 |
colorFrom: pink
|
|
@@ -8,18 +12,25 @@ sdk: docker
|
|
| 8 |
pinned: false
|
| 9 |
license: mit
|
| 10 |
short_description: The AI Courtroom That Exposes Bad Freelance Code
|
|
|
|
| 11 |
|
| 12 |
---
|
| 13 |
|
| 14 |
# Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
|
| 15 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 16 |
<div align="center">
|
| 17 |
|
| 18 |
# CodeTribunal
|
| 19 |
|
| 20 |
-
###
|
| 21 |
|
| 22 |
-
**
|
|
|
|
|
|
|
| 23 |
|
| 24 |
[](https://github.com/amineyagoub/CodeTribunal/actions/workflows/tests.yml)
|
| 25 |
[](https://www.python.org/downloads/)
|
|
@@ -30,18 +41,57 @@ short_description: The AI Courtroom That Exposes Bad Freelance Code
|
|
| 30 |
|
| 31 |
---
|
| 32 |
|
| 33 |
-
## The Problem
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 34 |
|
| 35 |
-
|
| 36 |
|
| 37 |
-
|
| 38 |
|
| 39 |
-
|
|
|
|
|
|
|
|
|
|
| 40 |
|
| 41 |
-
|
| 42 |
|
| 43 |
---
|
| 44 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 45 |
## How It Works
|
| 46 |
|
| 47 |
CodeTribunal runs a **6-phase pipeline**, each building on the last:
|
|
@@ -280,15 +330,6 @@ Test fixtures in `tests/fixtures/locale/` contain deliberately bad Python and Ja
|
|
| 280 |
|
| 281 |
---
|
| 282 |
|
| 283 |
-
## What Makes This a Strong Hackathon Entry
|
| 284 |
-
|
| 285 |
-
1. **System Complexity** β 6-phase pipeline with 8 agents, 4 custom tools, code graph, and streaming
|
| 286 |
-
2. **Effective Tool Use** β Agents use BaseTool with Pydantic schemas to read files, search patterns, and trace calls
|
| 287 |
-
3. **Context Handoffs** β CrewAI `context` parameter chains prosecution β defense β rebuttal
|
| 288 |
-
4. **Custom ReACT Engine** β Direct LiteLLM function calling with GLM-5 for reliable tool use
|
| 289 |
-
5. **Deterministic + AI** β GritQL provides ground-truth evidence, agents provide interpretation and debate
|
| 290 |
-
6. **Resilient** β Rate-limit retry, pipeline persistence, error recovery
|
| 291 |
-
|
| 292 |
---
|
| 293 |
|
| 294 |
<div align="center">
|
|
|
|
| 1 |
+
<<<<<<< HEAD
|
| 2 |
## <<<<<<< HEAD
|
| 3 |
|
| 4 |
+
=======
|
| 5 |
+
---
|
| 6 |
+
>>>>>>> cc6493d (docs: enhance README with frontmatter and add Hugging Face version)
|
| 7 |
title: CodeTribunal
|
| 8 |
emoji: π»
|
| 9 |
colorFrom: pink
|
|
|
|
| 12 |
pinned: false
|
| 13 |
license: mit
|
| 14 |
short_description: The AI Courtroom That Exposes Bad Freelance Code
|
| 15 |
+
<<<<<<< HEAD
|
| 16 |
|
| 17 |
---
|
| 18 |
|
| 19 |
# Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
|
| 20 |
|
| 21 |
+
=======
|
| 22 |
+
---
|
| 23 |
+
|
| 24 |
+
>>>>>>> cc6493d (docs: enhance README with frontmatter and add Hugging Face version)
|
| 25 |
<div align="center">
|
| 26 |
|
| 27 |
# CodeTribunal
|
| 28 |
|
| 29 |
+
### Put Freelance Code on Trial.
|
| 30 |
|
| 31 |
+
**Upload code. Get a verdict. Know the risk.**
|
| 32 |
+
|
| 33 |
+
Built with **GLM 5.1 + CrewAI + GritQL**
|
| 34 |
|
| 35 |
[](https://github.com/amineyagoub/CodeTribunal/actions/workflows/tests.yml)
|
| 36 |
[](https://www.python.org/downloads/)
|
|
|
|
| 41 |
|
| 42 |
---
|
| 43 |
|
| 44 |
+
## π¨ The Problem
|
| 45 |
+
|
| 46 |
+
Clients receive code they donβt understand.
|
| 47 |
+
|
| 48 |
+
- Looks clean⦠but hides security risks
|
| 49 |
+
- Passes linters⦠but fails in production
|
| 50 |
+
- Works⦠but is architecturally broken
|
| 51 |
+
|
| 52 |
+
**No one answers the only question that matters:**
|
| 53 |
+
|
| 54 |
+
> _Is this code safe, professional, and worth paying for?_
|
| 55 |
+
|
| 56 |
+
---
|
| 57 |
+
|
| 58 |
+
## The Solution
|
| 59 |
|
| 60 |
+
**CodeTribunal turns code review into a courtroom trial.**
|
| 61 |
|
| 62 |
+
Upload a `.zip` β get:
|
| 63 |
|
| 64 |
+
- Forensic evidence (AST-level)
|
| 65 |
+
- Multi-agent investigation
|
| 66 |
+
- AI courtroom debate
|
| 67 |
+
- Final verdict + risk score
|
| 68 |
|
| 69 |
+
> Not just analysis β **judgment**.
|
| 70 |
|
| 71 |
---
|
| 72 |
|
| 73 |
+
## π§ Why This Exist
|
| 74 |
+
|
| 75 |
+
### 1. Real System
|
| 76 |
+
|
| 77 |
+
- 6-phase pipeline
|
| 78 |
+
- 8 specialized agents
|
| 79 |
+
- Persistent execution engine
|
| 80 |
+
|
| 81 |
+
### 2. Agents That Actually Act
|
| 82 |
+
|
| 83 |
+
- File reads, pattern search, call tracing
|
| 84 |
+
- Real tool usage via function calling (not fake reasoning)
|
| 85 |
+
|
| 86 |
+
### 3. Deterministic + AI Hybrid
|
| 87 |
+
|
| 88 |
+
- **GritQL = ground truth**
|
| 89 |
+
- **Agents = interpretation + argument**
|
| 90 |
+
|
| 91 |
+
### 4. End-to-End Story
|
| 92 |
+
|
| 93 |
+
From raw code β evidence β debate β verdict β report
|
| 94 |
+
|
| 95 |
## How It Works
|
| 96 |
|
| 97 |
CodeTribunal runs a **6-phase pipeline**, each building on the last:
|
|
|
|
| 330 |
|
| 331 |
---
|
| 332 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 333 |
---
|
| 334 |
|
| 335 |
<div align="center">
|
README_HF.md
ADDED
|
@@ -0,0 +1,293 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
title: CodeTribunal
|
| 3 |
+
emoji: π»
|
| 4 |
+
colorFrom: pink
|
| 5 |
+
colorTo: red
|
| 6 |
+
sdk: docker
|
| 7 |
+
pinned: false
|
| 8 |
+
license: mit
|
| 9 |
+
short_description: The AI Courtroom That Exposes Bad Freelance Code
|
| 10 |
+
---
|
| 11 |
+
|
| 12 |
+
<div align="center">
|
| 13 |
+
|
| 14 |
+
# CodeTribunal
|
| 15 |
+
|
| 16 |
+
### Put Freelance Code on Trial.
|
| 17 |
+
|
| 18 |
+
**Upload code. Get a verdict. Know the risk.**
|
| 19 |
+
|
| 20 |
+
Built with **GLM 5.1 + CrewAI + GritQL**
|
| 21 |
+
|
| 22 |
+
[](https://github.com/amineyagoub/CodeTribunal/actions/workflows/tests.yml)
|
| 23 |
+
[](https://www.python.org/downloads/)
|
| 24 |
+
[](https://opensource.org/licenses/MIT)
|
| 25 |
+
[](https://build-with-glm-5-1-challenge.devpost.com)
|
| 26 |
+
|
| 27 |
+
</div>
|
| 28 |
+
|
| 29 |
+
---
|
| 30 |
+
|
| 31 |
+
## π¨ The Problem
|
| 32 |
+
|
| 33 |
+
Clients receive code they donβt understand.
|
| 34 |
+
|
| 35 |
+
- Looks clean⦠but hides security risks
|
| 36 |
+
- Passes linters⦠but fails in production
|
| 37 |
+
- Works⦠but is architecturally broken
|
| 38 |
+
|
| 39 |
+
**No one answers the only question that matters:**
|
| 40 |
+
|
| 41 |
+
> _Is this code safe, professional, and worth paying for?_
|
| 42 |
+
|
| 43 |
+
---
|
| 44 |
+
|
| 45 |
+
## The Solution
|
| 46 |
+
|
| 47 |
+
**CodeTribunal turns code review into a courtroom trial.**
|
| 48 |
+
|
| 49 |
+
Upload a `.zip` β get:
|
| 50 |
+
|
| 51 |
+
- Forensic evidence (AST-level)
|
| 52 |
+
- Multi-agent investigation
|
| 53 |
+
- AI courtroom debate
|
| 54 |
+
- Final verdict + risk score
|
| 55 |
+
|
| 56 |
+
> Not just analysis β **judgment**.
|
| 57 |
+
|
| 58 |
+
---
|
| 59 |
+
|
| 60 |
+
## π§ Why This Exist
|
| 61 |
+
|
| 62 |
+
### 1. Real System
|
| 63 |
+
|
| 64 |
+
- 6-phase pipeline
|
| 65 |
+
- 8 specialized agents
|
| 66 |
+
- Persistent execution engine
|
| 67 |
+
|
| 68 |
+
### 2. Agents That Actually Act
|
| 69 |
+
|
| 70 |
+
- File reads, pattern search, call tracing
|
| 71 |
+
- Real tool usage via function calling (not fake reasoning)
|
| 72 |
+
|
| 73 |
+
### 3. Deterministic + AI Hybrid
|
| 74 |
+
|
| 75 |
+
- **GritQL = ground truth**
|
| 76 |
+
- **Agents = interpretation + argument**
|
| 77 |
+
|
| 78 |
+
### 4. End-to-End Story
|
| 79 |
+
|
| 80 |
+
From raw code β evidence β debate β verdict β report
|
| 81 |
+
|
| 82 |
+
## How It Works
|
| 83 |
+
|
| 84 |
+
CodeTribunal runs a **6-phase pipeline**, each building on the last:
|
| 85 |
+
|
| 86 |
+
### Phase 1: Forensic Evidence (Deterministic β No LLM)
|
| 87 |
+
|
| 88 |
+
GritQL scans the entire codebase with **17 forensic patterns** across security and quality domains:
|
| 89 |
+
|
| 90 |
+
| Domain | Patterns | Examples |
|
| 91 |
+
| ----------- | -------- | ---------------------------------------------------------------------------------------- |
|
| 92 |
+
| π΄ Security | 13 | Hardcoded secrets, `eval()`, SQL injection, `pickle.load()`, `os.system()`, weak hashing |
|
| 93 |
+
| π‘ Quality | 4 | `TODO`, `FIXME`, `HACK` comments |
|
| 94 |
+
|
| 95 |
+
All scanning is **read-only** (`--dry-run`) and runs in **parallel** across patterns.
|
| 96 |
+
|
| 97 |
+
### Phase 2: Code Dependency Graph (AST β No LLM)
|
| 98 |
+
|
| 99 |
+
Python's `ast` module and regex-based JS parsing build a **lightweight dependency graph**:
|
| 100 |
+
|
| 101 |
+
- Nodes: files, functions, classes, imports
|
| 102 |
+
- Edges: calls, imports, containment, inheritance
|
| 103 |
+
- Enables call-chain tracing: `eval() β handle_request() β app.route()`
|
| 104 |
+
|
| 105 |
+
### Phase 3: Investigation (3 ReACT Agents + 4 Tools)
|
| 106 |
+
|
| 107 |
+
Three specialist investigators, each running a **genuine ReACT loop** (Reason β Act β Observe β Repeat) using **Z.ai's native function calling** via LiteLLM:
|
| 108 |
+
|
| 109 |
+
| Agent | Tools | Purpose |
|
| 110 |
+
| ------------------------- | --------------------------------------------------------- | ------------------------------------------ |
|
| 111 |
+
| Security Investigator | FileReader, PatternSearch, CodeGraphQuery, FindingContext | Find vulnerabilities, trace attack vectors |
|
| 112 |
+
| Quality Investigator | FileReader, FindingContext | Assess technical debt, detect negligence |
|
| 113 |
+
| Architecture Investigator | FileReader, CodeGraphQuery | Analyze structure, trace dependencies |
|
| 114 |
+
|
| 115 |
+
Each agent **autonomously decides which tools to call**, observes the results, and iterates. For example, the Security Investigator might:
|
| 116 |
+
|
| 117 |
+
1. Call `file_reader` to read a flagged file
|
| 118 |
+
2. Observe hardcoded secrets on specific lines
|
| 119 |
+
3. Call `code_graph_query` to trace where those secrets are used
|
| 120 |
+
4. Produce a detailed report with file paths, line numbers, and severity ratings
|
| 121 |
+
|
| 122 |
+
**Verified working**: GLM-5 + LiteLLM function calling confirmed. Agents make real tool calls that execute real code analysis.
|
| 123 |
+
|
| 124 |
+
### Phase 4: The Trial (3 Agents)
|
| 125 |
+
|
| 126 |
+
A courtroom debate between AI agents:
|
| 127 |
+
|
| 128 |
+
1. ** The Prosecutor** β builds the case for negligence, cites specific evidence
|
| 129 |
+
2. ** The Defense Attorney** β challenges claims, argues context and proportionality
|
| 130 |
+
3. ** Rebuttal** β the prosecutor responds to the defense
|
| 131 |
+
|
| 132 |
+
Agents use CrewAI's `context` parameter to chain arguments: prosecution output feeds into defense context, both feed into rebuttal.
|
| 133 |
+
|
| 134 |
+
### Phase 5: The Verdict
|
| 135 |
+
|
| 136 |
+
** The Judge** reviews all evidence, investigation reports, and the full trial transcript. Delivers:
|
| 137 |
+
|
| 138 |
+
- Overall ruling: GUILTY / MIXED / NOT GUILTY
|
| 139 |
+
- Reputational Risk Score (0-100)
|
| 140 |
+
- Findings summary with severity rankings
|
| 141 |
+
|
| 142 |
+
### Phase 6: Structured Report
|
| 143 |
+
|
| 144 |
+
** Verdict Report Agent** compiles everything into a professional report:
|
| 145 |
+
|
| 146 |
+
- Executive Summary
|
| 147 |
+
- Findings Table (sorted by severity)
|
| 148 |
+
- Per-Finding Analysis (impact, remediation, estimated fix effort)
|
| 149 |
+
- Sentencing Recommendations
|
| 150 |
+
|
| 151 |
+
---
|
| 152 |
+
|
| 153 |
+
## Architecture
|
| 154 |
+
|
| 155 |
+
```
|
| 156 |
+
ββββββββββββββββ
|
| 157 |
+
β Gradio UI β
|
| 158 |
+
β + Export β
|
| 159 |
+
ββββββββ¬ββββββββ
|
| 160 |
+
β
|
| 161 |
+
βββββββββββββΌβββββββββββββ
|
| 162 |
+
β Pipeline Engine β
|
| 163 |
+
β State Β· Persistence β
|
| 164 |
+
β Cancel Β· Resume β
|
| 165 |
+
βββββββββββββ¬βββββββββββββ
|
| 166 |
+
β
|
| 167 |
+
ββββββββββββ¬ββββββββββββΌββββββββββββ¬βββββββββββ
|
| 168 |
+
βΌ βΌ βΌ βΌ βΌ
|
| 169 |
+
βββββββββββ ββββββββ βββββββββββ βββββββββββ ββββββββ
|
| 170 |
+
βEvidence β βCode β βInvest. β β Trial β βReportβ
|
| 171 |
+
β Scanner β βGraph β β Agents β β Agents β βAgent β
|
| 172 |
+
β(GritQL) β β(AST) β β+ Tools β β β β β
|
| 173 |
+
βββββββββββ ββββββββ βββββββββββ βββββββββββ ββββββββ
|
| 174 |
+
β β β β β
|
| 175 |
+
ββββββββββββ΄ββββββββββββ΄ββββββββββββ΄βββββββββββ
|
| 176 |
+
β
|
| 177 |
+
βββββββββββββΌβββββββββββββ
|
| 178 |
+
β Custom Tool Layer β
|
| 179 |
+
β FileReader Β· Pattern β
|
| 180 |
+
β CodeGraph Β· Context β
|
| 181 |
+
ββββββββββββββββββββββββββ
|
| 182 |
+
```
|
| 183 |
+
|
| 184 |
+
### Key Design Decisions
|
| 185 |
+
|
| 186 |
+
| Decision | Why |
|
| 187 |
+
| ------------------------------------- | ------------------------------------------------------------------------------------------- |
|
| 188 |
+
| **Agents have tools, not text dumps** | Agents read files, search patterns, and trace calls on demand β scales to any codebase size |
|
| 189 |
+
| **ReACT loop via LiteLLM** | Direct function calling with GLM-5 β bypasses CrewAI's unreliable tool routing |
|
| 190 |
+
| **Pipeline state persisted to JSON** | Runs can resume after crashes. State is queryable |
|
| 191 |
+
| **GritQL for evidence** | AST-level pattern matching, not regex. Language-aware, precise |
|
| 192 |
+
| **Custom CrewAI tools (BaseTool)** | Pydantic-validated inputs, proper error handling, CrewAI-native integration |
|
| 193 |
+
| **Rate-limit retry with backoff** | Exponential backoff (4s β 64s) on Z.ai 429 errors β pipeline survives API spikes |
|
| 194 |
+
|
| 195 |
+
---
|
| 196 |
+
|
| 197 |
+
## Tech Stack
|
| 198 |
+
|
| 199 |
+
| Component | Technology | Purpose |
|
| 200 |
+
| -------------------- | -------------------------- | ---------------------------------------------------- |
|
| 201 |
+
| **LLM** | GLM 5.1 via Z.ai (LiteLLM) | Agent reasoning and debate |
|
| 202 |
+
| **Code Scanning** | GritQL | Deterministic AST-level pattern matching |
|
| 203 |
+
| **Multi-Agent** | CrewAI 1.12 | Agent orchestration, task chaining, context handoffs |
|
| 204 |
+
| **Function Calling** | LiteLLM | Direct ReACT loop with GLM-5 tool calling |
|
| 205 |
+
| **Code Graph** | Python `ast` + regex | Dependency graph (Python + JS) |
|
| 206 |
+
| **UI** | Gradio 6 | Streaming chatbot, file upload, export |
|
| 207 |
+
| **Export** | markdown-pdf (PyMuPDF) | PDF report generation from Markdown |
|
| 208 |
+
|
| 209 |
+
---
|
| 210 |
+
|
| 211 |
+
## Install
|
| 212 |
+
|
| 213 |
+
```bash
|
| 214 |
+
# Clone
|
| 215 |
+
git clone https://github.com/amineyagoub/CodeTribunal.git
|
| 216 |
+
cd CodeTribunal
|
| 217 |
+
|
| 218 |
+
# Install dependencies
|
| 219 |
+
pip install -e .
|
| 220 |
+
|
| 221 |
+
# Install GritQL CLI
|
| 222 |
+
npm install -g @getgrit/cli
|
| 223 |
+
|
| 224 |
+
# Configure
|
| 225 |
+
cp .env.example .env
|
| 226 |
+
# Edit .env: set ZAI_API_KEY (get one at https://open.bigmodel.cn/)
|
| 227 |
+
```
|
| 228 |
+
|
| 229 |
+
### Requirements
|
| 230 |
+
|
| 231 |
+
- Python 3.11+
|
| 232 |
+
- Node.js (for GritQL CLI)
|
| 233 |
+
- Z.ai API key ([get one here](https://open.bigmodel.cn/))
|
| 234 |
+
|
| 235 |
+
---
|
| 236 |
+
|
| 237 |
+
## Usage
|
| 238 |
+
|
| 239 |
+
### Web UI (Recommended)
|
| 240 |
+
|
| 241 |
+
```bash
|
| 242 |
+
python3 -m code_tribunal.app
|
| 243 |
+
```
|
| 244 |
+
|
| 245 |
+
Open http://localhost:7860, upload a `.zip` of code, and watch the trial unfold.
|
| 246 |
+
|
| 247 |
+
---
|
| 248 |
+
|
| 249 |
+
## Features
|
| 250 |
+
|
| 251 |
+
| Feature | Details |
|
| 252 |
+
| ------------------------------ | --------------------------------------------------------------------------------------- |
|
| 253 |
+
| **4 Custom Tools** | FileReader, PatternSearch, CodeGraphQuery, FindingContext β agents actively investigate |
|
| 254 |
+
| **8 Specialized Agents** | 3 investigators, prosecutor, defense, rebuttal, judge, verdict report, expert witness |
|
| 255 |
+
| **ReACT Engine** | Custom Reason-Act-Observe loop via LiteLLM function calling with GLM-5 |
|
| 256 |
+
| **Code Dependency Graph** | AST-based (Python + JS), with call-chain tracing and impact analysis |
|
| 257 |
+
| **Parallel Evidence Scanning** | ThreadPoolExecutor for GritQL patterns β 4x faster than sequential |
|
| 258 |
+
| **Rate-Limit Resilience** | Exponential backoff retry on 429 errors β survives API rate limits |
|
| 259 |
+
| **Pipeline Persistence** | State saved to JSON, runs can resume after interruption |
|
| 260 |
+
| **Deduplication** | Same file+line merged into one finding with multiple categories |
|
| 261 |
+
| **Zip Safety** | Zip-slip attack prevention |
|
| 262 |
+
| **Streaming UI** | Real-time pipeline progress in Gradio Chatbot with phase indicators |
|
| 263 |
+
| **Export** | Markdown and PDF report generation |
|
| 264 |
+
|
| 265 |
+
---
|
| 266 |
+
|
| 267 |
+
## π§ͺ Testing
|
| 268 |
+
|
| 269 |
+
```bash
|
| 270 |
+
# Run evidence scan on test fixtures
|
| 271 |
+
code-tribunal tests/fixtures/locale/ --evidence-only
|
| 272 |
+
|
| 273 |
+
# Run Python tests
|
| 274 |
+
pytest tests/
|
| 275 |
+
```
|
| 276 |
+
|
| 277 |
+
Test fixtures in `tests/fixtures/locale/` contain deliberately bad Python and JavaScript code with:
|
| 278 |
+
|
| 279 |
+
- Hardcoded passwords, API keys, AWS secrets, Stripe keys, JWT secrets
|
| 280 |
+
- SQL injection via f-strings and template literals
|
| 281 |
+
- `eval()`, `pickle.load()`, `os.system()`, `subprocess.call(shell=True)`
|
| 282 |
+
- MD5 hashing
|
| 283 |
+
- TODO, FIXME, HACK comments
|
| 284 |
+
|
| 285 |
+
---
|
| 286 |
+
|
| 287 |
+
---
|
| 288 |
+
|
| 289 |
+
<div align="center">
|
| 290 |
+
|
| 291 |
+
Built for the [Build with GLM 5.1](https://build-with-glm-5-1-challenge.devpost.com) hackathon.
|
| 292 |
+
|
| 293 |
+
</div>
|