HumanEval: Benchmark Detailed Guide
Detailed guide to HumanEval: category, metrics, sources and applicable models.
Overview
OpenAI 发布的 164 道 Python 编程任务,每题包含函数签名、文档字符串、函数体与单元测试,评估模型的代码生成能力(pass@1)。
Metrics
| Metric | Unit | Direction |
|---|---|---|
| pass@1 | pass@1 | ↑ Higher is better |
Sources
Model Score Ranking
HumanEval
Descrizione
OpenAI 发布的 164 道 Python 编程任务,每题包含函数签名、文档字符串、函数体与单元测试,评估模型的代码生成能力(pass@1)。
Specifiche principali
| Categoria | Licenza | Ultimo aggiornamento |
|---|---|---|
| coding | MIT | 2022-01-01 |
Benchmark
| Unità |
|---|
| pass@1 |