Songyang Zhang
|
98435dd98e
|
[Feature] Update o1 evaluation with JudgeLLM (#1795)
* Update Generic LLM Evaluator
* Update o1 style evaluator
|
2024-12-30 17:31:00 +08:00 |
|
bittersweet1999
|
38dba9919b
|
[Fix] Fix Subjective summarizer order error (#1767)
* fix pip version
* fix pip version
* fix order error
|
2024-12-18 13:21:31 +08:00 |
|
bittersweet1999
|
08d63b5bf3
|
[Fix] Fix error in subjective default summarizer (#1740)
* fix pip version
* fix pip version
* fix summarizer bug
|
2024-12-06 11:03:53 +08:00 |
|
Haoran Que
|
4fe251729b
|
Upload HelloBench (#1607)
* upload hellobench
* update hellobench
* update readme.md
* update eval_hellobench.py
* update lastest
---------
Co-authored-by: bittersweet1999 <148421775+bittersweet1999@users.noreply.github.com>
|
2024-10-15 17:11:37 +08:00 |
|
bittersweet1999
|
fa54aa62f6
|
[Feature] Add Judgerbench and reorg subeval (#1593)
* fix pip version
* fix pip version
* update (#1522)
Co-authored-by: zhulin1 <zhulin1@pjlab.org.cn>
* [Feature] Update Models (#1518)
* Update Models
* Update
* Update humanevalx
* Update
* Update
* [Feature] Dataset prompts update for ARC, BoolQ, Race (#1527)
add judgerbench and reorg sub
add judgerbench and reorg subeval
add judgerbench and reorg subeval
* add judgerbench and reorg subeval
* add judgerbench and reorg subeval
* add judgerbench and reorg subeval
* add judgerbench and reorg subeval
---------
Co-authored-by: zhulinJulia24 <145004780+zhulinJulia24@users.noreply.github.com>
Co-authored-by: zhulin1 <zhulin1@pjlab.org.cn>
Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com>
Co-authored-by: Linchen Xiao <xxllcc1993@gmail.com>
|
2024-10-15 16:36:05 +08:00 |
|