Benchmarks¶
All benchmarks in Skills Tree are reproducible: they include methodology, datasets, and test scripts.
Available Benchmarks¶
Reasoning¶
| Benchmark | Dataset | Winner | Margin | Link |
|---|---|---|---|---|
| ReAct vs LATS | HotpotQA | LATS | +8.3% accuracy | View |
Memory & Retrieval¶
| Benchmark | Dataset | Winner | Margin | Link |
|---|---|---|---|---|
| RAG retrieval strategies | Custom | HyDE | +12% recall | View |
| Memory injection methods | Custom | Top-K semantic | Best cost/quality | View |
Tool Use¶
| Benchmark | Dataset | Winner | Margin | Link |
|---|---|---|---|---|
| Function calling | ToolBench | Claude 3.7 | +6% accuracy | View |
Reproducing a Benchmark¶
git clone https://github.com/SamoTech/skills-tree.git
cd skills-tree
pip install -r requirements.txt
python benchmarks/reasoning/react-vs-lats.py
Contributing Benchmarks¶
Benchmarks are among the highest-value contributions. To add one:
- Create
benchmarks/{category}/{name}.mdfollowing the existing format - Include: dataset, methodology, results table, reproduction script
- Open a PR with title format:
benchmark: [skill-a] vs [skill-b]
See contributing guide for full details.