OpenCompass

mirror of https://github.com/open-compass/opencompass.git synced 2025-05-30 16:03:24 +08:00

Author	SHA1	Message	Date
Mo Li	f2af49337d	[Feature] Add ATC Choice Version (#1019 ) * Squashed commit of the following: commit c48ad194c3976dc63d1b60d8c8ab2d5ff9e1cbfe Author: DseidLi <2568818204@qq.com> Date: Tue Apr 2 16:57:43 2024 +0800 add atc_choice commit 3ac6efea29619573e6fac8fa3cce464853dcead0 Merge: `2d4e559` 8e3a9c3 Author: DseidLi <2568818204@qq.com> Date: Tue Apr 2 16:41:38 2024 +0800 Merge branch 'atc_choice' into atc_add_choice commit 8e3a9c396a3e5546d3faf584183f6fd60b974d5e Merge: 150a036 `0a6a03f` Author: DseidLi <2568818204@qq.com> Date: Tue Mar 26 04:47:07 2024 +0800 Merge branch 'main' into atc_choice Conflicts: configs/summarizers/needlebench.py opencompass/datasets/needlebench/multi.py opencompass/datasets/needlebench/origin.py opencompass/datasets/needlebench/parallel.py commit 150a036d6d990f26a57c974d1af83d88c31a0f9d Merge: 8d6ac9a 940dd18 Author: DseidLi <2568818204@qq.com> Date: Wed Mar 20 03:49:08 2024 +0800 Merge branch 'needlebench_fix' into atc_choice commit 8d6ac9a1a43b1c9d0f0ea27e7d58968a203ea898 Author: DseidLi <2568818204@qq.com> Date: Wed Mar 20 03:41:49 2024 +0800 optimize needlebench code commit 940dd18a4270f24bc69edd2a780182c68918e1a9 Author: DseidLi <2568818204@qq.com> Date: Wed Mar 20 03:39:46 2024 +0800 fix vllm commit d8be6877bc41051f3edcc0421c462c834c0f1c9a Merge: ecad78a `2527fda` Author: DseidLi <2568818204@qq.com> Date: Tue Mar 19 21:07:08 2024 +0800 Merge remote-tracking branch 'origin/add_1M_dataset' into atc_choice commit `2527fda8a5` Author: DseidLi <2568818204@qq.com> Date: Tue Mar 19 16:03:40 2024 +0800 add model configs commit `75425acdf8` Author: DseidLi <2568818204@qq.com> Date: Tue Mar 19 16:02:15 2024 +0800 add prompt postion args commit `367ba1ba61` Author: DseidLi <2568818204@qq.com> Date: Wed Feb 28 21:40:00 2024 +0800 add Needlebench-1000K configs commit ecad78af14c4bb00fe325779114b384c57ab30bf Author: DseidLi <2568818204@qq.com> Date: Thu Mar 14 22:08:32 2024 +0800 fix atc commit 08772c0787b18872abadc9ffec3223941a5ee0c2 Merge: 9f3f8cf `caf1cf8` Author: DseidLi <2568818204@qq.com> Date: Thu Mar 14 22:07:28 2024 +0800 Merge branch 'main' into atc_choice Conflicts: configs/datasets/needlebench/readme.md configs/datasets/needlebench/readme_zh-CN.md configs/summarizers/needlebench.py opencompass/datasets/needlebench/atc.py opencompass/summarizers/needlebench.py commit 9f3f8cfb4452722734d334114ac1d14110e57406 Author: DseidLi <2568818204@qq.com> Date: Thu Mar 14 21:35:53 2024 +0800 add atc-choice test commit 52be7c1202376b4e09821188b826f1a805328129 Author: DseidLi <2568818204@qq.com> Date: Wed Mar 6 02:54:15 2024 +0800 update needlebench randomseed and add vllm qwen14b commit fc1effce596ae2e5ece4933e8cd34aef8e64a6f9 Merge: 4e747ed `caf1cf8` Author: DseidLi <2568818204@qq.com> Date: Wed Mar 6 02:51:14 2024 +0800 Merge branch 'main' into add_model_configs commit 31834f9b23af3354ac3581ec86d693d0f05cdd1c Merge: 7dabc82 `120bf8b` Author: DseidLi <2568818204@qq.com> Date: Sun Mar 3 23:29:42 2024 +0800 Merge branch 'main' of https://github.com/open-compass/opencompass into atc_choice commit 4e747ed1988ddbcfcc7fff334601259ade72d363 Author: DseidLi <2568818204@qq.com> Date: Sun Mar 3 22:15:25 2024 +0800 add internlm2-lmdeploy model and gemma configs commit 7dabc828123d711c8cf834d6aab4137bb55e85ed Author: DseidLi <2568818204@qq.com> Date: Sat Mar 2 17:26:15 2024 +0800 add atc choice version -ZH commit `996f8ae43d` Author: DseidLi <2568818204@qq.com> Date: Wed Feb 28 16:58:56 2024 +0800 update readme for needlebench commit `f7266e873c` Author: DseidLi <2568818204@qq.com> Date: Wed Feb 28 16:44:53 2024 +0800 move readme.md commit `1c7375681d` Author: DseidLi <2568818204@qq.com> Date: Wed Feb 28 16:38:31 2024 +0800 fix linting error commit `b6524f3ebf` Author: DseidLi <2568818204@qq.com> Date: Wed Feb 28 16:33:51 2024 +0800 lint summarizer commit `c0d1190e39` Author: DseidLi <2568818204@qq.com> Date: Wed Feb 28 16:29:03 2024 +0800 add needlebench intro, fix summarizer commit `0965baf785` Author: DseidLi <2568818204@qq.com> Date: Mon Feb 26 13:31:26 2024 +0800 fix bug in needlebench summarizer commit `5d32b31eb8` Author: DseidLi <2568818204@qq.com> Date: Sat Feb 24 03:19:08 2024 +0800 update act prompt commit `af82a7f085` Merge: `32bf9fe` `53fe788` Author: DseidLi <2568818204@qq.com> Date: Fri Feb 23 17:50:32 2024 +0800 Merge remote-tracking branch 'upstream/main' into needlebench commit `32bf9fe802` Author: DseidLi <2568818204@qq.com> Date: Fri Feb 23 17:31:32 2024 +0800 simplify needlebench 32k, 128k, 200k for eval commit `a7cb025e05` Author: DseidLi <2568818204@qq.com> Date: Fri Feb 23 14:48:58 2024 +0800 add needlebench * fix summarizer * remove repeated code * remove chinese comments	2024-04-07 15:46:20 +08:00
Mo Li	b50d163265	[Fix] Refactor Needlebench Configs for CLI Testing Support (#1020 ) * add needlebench datasets suffix * fix import * update run.py args for summarizer key and dataset suffix * update utils/run.py	2024-04-07 15:12:56 +08:00
bittersweet1999	2d4e559763	[Feature] Add multi-model judge and fix some problems (#1016 ) * support multi-model judge and moe judge * test_moe * test_moe * test * add moe judge * support multi-judge-model	2024-04-02 11:52:06 +08:00
bittersweet1999	02e7eec911	[Feature] Support AlpacaEval_V2 (#1006 ) * support alpacaeval_v2 * support alpacaeval * update docs * update docs	2024-03-28 16:49:04 +08:00
Mo Li	0a6a03fe1a	[Feature] update needlebench and configs (#986 ) * add Needlebench-1000K configs * add prompt postion args * add model configs * Update parallel.py * fix lint	2024-03-25 18:05:01 +08:00
Chaseldot	1d3198554b	[Fix] base.py change status into list (#994 )	2024-03-22 17:06:34 +08:00
Ke Bao	e415ddf96a	[Fix] Fix turbomind_tis (#992 )	2024-03-22 15:50:12 +08:00
Connor-Shen	0221d30877	[Fix] Update APPS/TACO (#988 ) * [Feature] update apps/taco * [Feature] update apps/taco	2024-03-19 20:21:39 +08:00
Connor-Shen	8a3c6e51ed	[Feature] Update APPS (#985 ) * update post process * update post process	2024-03-19 15:47:05 +08:00
Connor-Shen	d92595b671	[Feat] Support TACO (#966 ) * [Feat] Support TACO * update README * update README	2024-03-19 15:39:16 +08:00
bittersweet1999	c78a4df923	add support for set prediction path (#984 )	2024-03-19 14:32:15 +08:00
Jingming	89a8a8917b	[Feature] Add the implement of QuALITY datasets (#976 ) #976	2024-03-15 21:22:38 +08:00
Connor-Shen	3098d78845	[Bench] Support APPS (#963 ) * [Feat] support apps * [Feat] support apps * [Feat] support apps * update README	2024-03-13 16:09:23 +08:00
Fengzhe Zhou	ab6cdb2be8	[Sync] Bump version 0.2.3 (#957 )	2024-03-12 11:51:56 +08:00
Fengzhe Zhou	64fde73b15	[Fix] Use logger.error on failure (#960 )	2024-03-12 11:51:39 +08:00
Fengzhe Zhou	bdd85358cc	[Sync] update 20240308 (#953 )	2024-03-11 22:34:19 +08:00
bittersweet1999	848e7c8a76	[fix] add different temp for different question in mtbench (#954 ) * add temp for mtbench * add document for mtbench * add document for mtbench	2024-03-11 17:24:39 +08:00
Yang Yong	3829be87b1	Fix LightllmApi ppl test (#951 )	2024-03-08 12:04:44 +08:00
Yang Yong	107e022cf4	Support prompt template for LightllmApi. Update LightllmApi token bucket. (#945 )	2024-03-06 15:33:53 +08:00
RunningLeon	c54a5d3b0f	Support get_ppl for TurbomindModel (#878 ) * update ppl for turbomindmodel * update api_server * rename config and set thread_safe for pytorch engine if possible	2024-03-06 11:44:19 +08:00
Fengzhe Zhou	b03d5dc531	[Sync] Sync Internal (#941 )	2024-03-04 14:42:36 +08:00
yuantao2108	bbec7d8733	[Feature] add lveval benchmark (#914 ) * add lveval benchmark * add LVEval readme file * update LVEval readme file * Update configs/eval_bluelm_32k_lveval.py * Update configs/eval_llama2_7b_lveval.py --------- Co-authored-by: yuantao <yuantao@infini-ai.com> Co-authored-by: Mo Li <82895469+DseidLi@users.noreply.github.com>	2024-03-04 11:22:03 +08:00
Mo Li	8142f399a8	[Feature] Upgrade the needle-in-a-haystack experiment to Needlebench (#913 ) * add needlebench * simplify needlebench 32k, 128k, 200k for eval * update act prompt * fix bug in needlebench summarizer * add needlebench intro, fix summarizer * lint summarizer * fix linting error * move readme.md * update readme for needlebench * update docs of needlebench * simplify needlebench summarizers	2024-03-04 11:10:52 +08:00
Kdump	3e9844ed33	[Fix]Fixed the problem of never entering task.run() mode in local scheduling mode. (#930 ) * Fixed the problem of never entering task.run() mode in local scheduling mode. get_command_template方法中为命令行前缀添加了CUDA_VISIBLE_DEVICES=或set CUDA_VISIBLE_DEVICES=。导致task.run()分支失效。 --------- CUDA_VISIBLE_DEVICES= or set CUDA_VISIBLE_DEVICES= is added to the command line prefix in the get_command_template method. Causes the task.run() branch to fail. * [Fix]Fixed the problem of never entering task.run() mode in local scheduling mode. get_command_template方法中为命令行前缀添加了CUDA_VISIBLE_DEVICES=或set CUDA_VISIBLE_DEVICES=。导致task.run()分支失效。 --- CUDA_VISIBLE_DEVICES= or set CUDA_VISIBLE_DEVICES= is added to the command line prefix in the get_command_template method. Causes the task.run() branch to fail. * [Fix]Fixed the problem of never entering task.run() mode in local scheduling mode. get_command_template方法中为命令行前缀添加了CUDA_VISIBLE_DEVICES=或set CUDA_VISIBLE_DEVICES=。导致task.run()分支失效。 CUDA_VISIBLE_DEVICES= or set CUDA_VISIBLE_DEVICES= is added to the command line prefix in the get_command_template method. Causes the task.run() branch to fail.	2024-02-29 14:35:45 +08:00
Skyfall-xzz	4c45a71bbc	[Feature] Support OpenFinData (#896 ) * [Feature] Support OpenFinData * add README for OpenFinData * update README	2024-02-29 12:55:07 +08:00
bittersweet1999	001e77fea2	[Feature] add support for gemini (#931 ) * add gemini * add gemini * add gemini	2024-02-28 19:38:34 +08:00
Fengzhe Zhou	9afbfa3639	[Sync] Fix TEvalEvaluator (#929 )	2024-02-28 16:05:30 +08:00
Fengzhe Zhou	5ce8e0450e	[Fix] Fix type hint in IFEval (#915 )	2024-02-28 10:53:40 +08:00
Jingming	53fe788d27	[Fix] fix ifeval (#909 )	2024-02-23 16:52:03 +08:00
bittersweet1999	45c606bcd0	[Fix] Fix IFEval (#906 ) * fix ifeval * fix ifeval * fix ifeval * fix ifeval	2024-02-22 16:51:34 +08:00
RunningLeon	32ba0b074e	Support lmdeploy pytorch engine (#875 ) * add lmdeploy pytorch model * fix * speed up encoding and decoding * fix * change tokenizer	2024-02-22 03:46:07 -03:00
Yang Yong	b6e21ece38	Support LightllmApi input_format (#888 )	2024-02-19 10:02:59 +08:00
Fengzhe Zhou	08133e060a	[Sync] Bump version to 0.2.2 (#880 )	2024-02-07 10:45:48 +08:00
hailsham	dd444685bb	fix bug of gsm8k_postprocess (#863 ) * fix bug of gsm8k_postprocess * update postprocess --------- Co-authored-by: Lei Fei <SENSETIME\leifei1@cn3114002087l.domain.sensetime.com> Co-authored-by: Leymore <zfz-960727@163.com>	2024-02-06 23:52:47 +08:00
Connor-Shen	444d8d9507	[feat] support multipl-e (#846 ) * [feat] support humaneval_multipl-e * format --------- Co-authored-by: Leymore <zfz-960727@163.com>	2024-02-06 23:30:28 +08:00
Yggdrasill7D6	a6c49f15ce	fix lawbench 2-1 f0.5 score calculation bug (#795 ) * fix lawbench 2-1 f0.5 score calculation bug * use path in overall datasets folder --------- Co-authored-by: Leymore <zfz-960727@163.com>	2024-02-06 22:20:11 +08:00
bittersweet1999	1c8e193de8	[Fix] hotfix for mtbench (#877 ) * hotfix for mtbench * hotfix	2024-02-06 21:26:47 +08:00
Fengzhe Zhou	d34ba11106	[Sync] Merge branch 'dev' into zfz/update-keyset-demo (#876 )	2024-02-05 23:29:10 +08:00
Skyfall-xzz	7ad1168062	Support NPHardEval (#835 ) * support NPHardEval * add .md file and fix minor bugs * refactor and minor fix --------- Co-authored-by: Leymore <zfz-960727@163.com>	2024-02-05 15:52:28 +08:00
Yuchen Yan	fed7d800c6	[Fix] Fix error in gsm8k evaluator (#782 ) Co-authored-by: jiangjin1999 <1261842974@qq.com>	2024-02-04 22:55:11 +08:00
bittersweet1999	7806cd0f64	[Feature] support alpacaeval (#809 ) * support alpacaeval_v1 * Update opencompass/summarizers/subjective/__init__.py Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com> * Update opencompass/summarizers/subjective/alpacaeval_v1.py Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com> * fix conflict * support alpacaeval v2 * support alpacav2 --------- Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com>	2024-02-04 14:18:36 +08:00
RunningLeon	4c87e777d8	[Feature] Add end_str for turbomind (#859 ) * fix * update * fix internlm1 * fix docs * remove sys	2024-02-01 22:31:14 +08:00
bittersweet1999	5c6dc908cd	fix compass arena (#854 )	2024-01-30 16:34:38 +08:00
Songyang Zhang	cdca59ff49	[Fix] Update Zhipu API and Fix issue min_out_len issue of API models (#847 ) * Update zhipu api and fix min_out_len issue of API class * Update example * Update example	2024-01-28 14:52:43 +08:00
Jingming	2801883351	[Fix] Fix acc of IFEval (#849 ) * [Feature] Add IFEval * [Fix] Changing the Score Rule.	2024-01-27 22:27:07 +08:00
Xiaoming Shi	35aace776a	[Fix] Update MedBench (#845 )	2024-01-26 17:56:13 +08:00
Songyang Zhang	8ed022b4c4	Update Sensetime API (#844 )	2024-01-26 16:40:49 +08:00
Hubert	4aa74565e2	[Feat] minor update agent related (#839 ) * [Feat] update cibench * [Feat] Support CIBench * [Feat] Support CIBench * [Feat] Support CIBench * [Feat] Support CIBench	2024-01-26 14:15:51 +08:00
Fengzhe Zhou	0991dd33a0	[Sync] Updata dataset cfg for internMath (#837 ) Co-authored-by: liuhongwei <liuhongwei@pjlab.org.cn>	2024-01-24 16:30:32 +08:00
Songyang Zhang	793e32c9cc	[Feature] Update API implementation (#834 )	2024-01-24 13:35:21 +08:00
bittersweet1999	2ee8e8a1a1	[Feature] add mtbench (#829 ) * add mtbench * add mtbench * Update configs/datasets/subjective/multiround/mtbench_judgeby_gpt4.py Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com> * Update configs/datasets/subjective/multiround/mtbench_judgeby_gpt4.py Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com> * Update opencompass/datasets/subjective/__init__.py Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com> * Update opencompass/datasets/subjective/mtbench.py Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com> * fix mtbench --------- Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com>	2024-01-24 12:11:47 +08:00
Jingming	e059a5c2bf	[Feature] Add IFEval (#813 ) * [Feature] Add IFEval * [Doc] add introduction of IFEval	2024-01-23 20:07:49 +08:00
bittersweet1999	3d9bb4aed7	[Fix] fix strings (#833 ) * add compass arena * add compass_arena * add compass arena * Update opencompass/summarizers/subjective/compass_arena.py Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com> * Update opencompass/summarizers/subjective/__init__.py Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com> * Update opencompass/datasets/subjective/compass_arena.py Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com> * Update opencompass/datasets/subjective/__init__.py Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com> * Update configs/eval_subjective_compassarena.py Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com> * Update configs/datasets/subjective/compassarena/compassarena_compare.py Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com> * Update configs/eval_subjective_compassarena.py Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com> * Update configs/datasets/subjective/compassarena/compassarena_compare.py Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com> * fix check position bias * fix string --------- Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com>	2024-01-23 10:57:26 +00:00
bittersweet1999	2d4da8dd02	[Feature] Add CompassArena (#828 ) * add compass arena * add compass_arena * add compass arena * Update opencompass/summarizers/subjective/compass_arena.py Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com> * Update opencompass/summarizers/subjective/__init__.py Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com> * Update opencompass/datasets/subjective/compass_arena.py Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com> * Update opencompass/datasets/subjective/__init__.py Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com> * Update configs/eval_subjective_compassarena.py Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com> * Update configs/datasets/subjective/compassarena/compassarena_compare.py Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com> * Update configs/eval_subjective_compassarena.py Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com> * Update configs/datasets/subjective/compassarena/compassarena_compare.py Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com> * fix check position bias --------- Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com>	2024-01-23 15:12:46 +08:00
Guo Qipeng	e975a96fa1	Update cdme config and evaluator (#812 ) * update cdme config and evaluator * fix cdme prompt * move CDME trim post-processor as a separate evaluator --------- Co-authored-by: 郭琦鹏 <guoqipeng@pjlab.org.cn>	2024-01-19 11:29:27 +08:00
Yang Yong	f09a2ff418	Add LightllmApi KeyError log & Update doc (#816 ) * Add LightllmApi KeyError log * Update LightllmApi doc	2024-01-18 22:23:38 +08:00
RunningLeon	61fe873c89	[Fix] Fix turbomind and update docs (#808 ) * update * update docs * add engine_config and gen_config in eval_config * update * fix * fix * fix * fix docstr * fix url	2024-01-18 14:41:35 +08:00
Fengzhe Zhou	b4afe3e7c1	[Sync] Add InternLM2 Keyset Evaluation Demo (#807 ) Co-authored-by: zhangyifan1 <zhangyifan1@pjlab.org.cn>	2024-01-17 13:48:12 +08:00
Mo Li	acae560911	Added support for multi-needle testing in needle-in-a-haystack test (#802 ) * Add NeedleInAHaystack Test * Apply pre-commit formatting * Update configs/eval_hf_internlm_chat_20b_cdme.py Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com> * add needle in haystack test * update needle in haystack test * update plot function in tools_needleinahaystack.py * optimizing needleinahaystack dataset generation strategy * modify minor formatting issues * add English version support * change NeedleInAHaystackDataset to dynamic loading * change NeedleInAHaystackDataset to dynamic loading * fix needleinahaystack test eval bug * fix needleinahaystack config bug * Added support for multi-needle testing in needle-in-a-haystack test * Optimize the code for plotting in the needle-in-a-haystack test. * Correct the typo in the dataset parameters. * update needleinahaystack test docs --------- Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com>	2024-01-17 13:47:34 +08:00
RunningLeon	0836aec67b	[Feature] Update evaluate turbomind (#804 ) * update * fix * fix * fix	2024-01-17 11:09:50 +08:00
bittersweet1999	814b3f73bd	reorganize subject files (#801 )	2024-01-16 18:03:11 +08:00
bittersweet1999	83d6c48378	[Feature] Add configs for creationbench (#791 ) * add creationv2_zh * add creationv2_zh * add eng config for creationbench * add eng config for creationbench * add eng config for creationbench	2024-01-12 14:20:21 +08:00
notoschord	d3a0ddc3ef	[Feature] Add support for Nanbeige API (#786 ) Co-authored-by: notoschord <wangzekai@kanzhun.com>	2024-01-11 13:54:27 +08:00
bittersweet1999	5679edb490	add temperature in alles (#787 )	2024-01-11 03:57:24 +00:00
Xiaoming Shi	ad872a5dc2	[Feature] Update MedBench (#779 ) * update medbench * medbench update * format medbench * format * Update * update * update * update suffix --------- Co-authored-by: 施晓明 <PJLAB\shixiaoming@pjnl104220118l.pjlab.org> Co-authored-by: Leymore <zfz-960727@163.com>	2024-01-09 11:42:44 +08:00
Fengzhe Zhou	a74e4c1a8d	[Sync] Bump version to 0.2.1 (#778 )	2024-01-08 14:56:28 +00:00
Fengzhe Zhou	32f40a8f83	[Sync] Sync with internal codes 2023.01.08 (#777 )	2024-01-08 14:07:24 +00:00
jiangjin1999	8194199d79	[Feature] _batch_generate function, add the MultiTokenEOSCriteria (#772 ) * jiangjin1999: in the _batch_generate function, add the MultiTokenEOSCriteria feature to speed up inference. * jiangjin1999: in the _batch_generate function, add the MultiTokenEOSCriteria feature to speed up inference. --------- Co-authored-by: jiangjin08 <jiangjin08@MBP-2F32S5MD6P-0029.local> Co-authored-by: jiangjin08 <jiangjin08@a.sh.vip.dianping.com>	2024-01-08 16:40:02 +08:00
liyucheng09	0b2863039e	[Feature] Contamination analysis for MMLU, Hellaswag, and ARC_c (#699 ) * Contamination analysis for ARC_c, mmlu, and Hellaswag * update `eval_contamination.py` * update `contamination.py` summarizer * fix `eval_contamination.py` * add mmlu groups for contamination analysis	2024-01-08 15:51:48 +08:00
Connor-Shen	30a90d8dd8	Support Mbpp_plus dataset (#770 ) * support mbpp+ * support mbpp+ * minor fix * [Feat] minor fix --------- Co-authored-by: yingfhu <yingfhu@gmail.com>	2024-01-05 22:01:57 +08:00
bittersweet1999	3c606cb712	quick fix for postprocess pred extraction (#771 )	2024-01-05 21:10:18 +08:00
bittersweet1999	2163f9398f	[Feature] add subject ir dataset (#755 ) * add subject ir * Add ir dataset * Add ir dataset	2024-01-05 12:00:57 +00:00
bittersweet1999	be369c3e06	[Feature] Add multi_round dataset evaluation (#766 ) * multi_round dataset * add multi_round evaluation	2024-01-04 10:37:52 +00:00
bittersweet1999	7cd65d49d8	[Fix] Fix small bug in alignbench (#764 ) * fix small bugs * fix small bugs	2024-01-03 07:44:53 +00:00
Chris Liu	3eb225a5e6	[Feature] Support LLaMA2-Accessory (#732 ) * Support LLaMA2-Accessory * remove strip * clear imports * reformat * fix lint * fix lint * update readme * update readme * update readme * update readme	2024-01-02 20:48:51 +08:00
HUANG Fei	ba027eeeac	[Feature] Add support of qwen api (#735 )	2024-01-02 20:47:12 +08:00
Mo Li	33f8df1ca3	[Update] Change NeedleInAHaystackDataset to dynamic dataset loading (#754 ) * Add NeedleInAHaystack Test * Apply pre-commit formatting * Update configs/eval_hf_internlm_chat_20b_cdme.py Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com> * add needle in haystack test * update needle in haystack test * update plot function in tools_needleinahaystack.py * optimizing needleinahaystack dataset generation strategy * modify minor formatting issues * add English version support * change NeedleInAHaystackDataset to dynamic loading * change NeedleInAHaystackDataset to dynamic loading * fix needleinahaystack test eval bug * fix needleinahaystack config bug --------- Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com>	2024-01-02 17:22:56 +08:00
Francis-llgg	b69fe2343b	[Feature] Add GPQA Dataset (#729 ) * check * message * add * change prompt * change a para nameq * modify name of the file * delete an useless file	2024-01-01 15:54:40 +08:00
Francis-llgg	ef3ae63539	[Feature] Add new dataset mastermath2024v1 (#744 ) * add new dataset mastermath2024v1 * change it to simplified chinese prompt * change file name	2024-01-01 15:53:24 +08:00
Mo Li	17b8e929dd	[Feature] Update plot function in tools_needleinahaystack.py (#747 ) * Add NeedleInAHaystack Test * Apply pre-commit formatting * Update configs/eval_hf_internlm_chat_20b_cdme.py Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com> * add needle in haystack test * update needle in haystack test * update plot function in tools_needleinahaystack.py * optimizing needleinahaystack dataset generation strategy * modify minor formatting issues --------- Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com>	2023-12-29 18:51:09 +08:00
Hubert	327951087f	[Feat] update code config (#749 ) * [Feat] update code dataset * [Feat] update code dataset * [Feat] update code dataset	2023-12-29 18:46:34 +08:00
bittersweet1999	fe0b717033	add creationbench (#753 )	2023-12-29 10:03:44 +00:00
Connor-Shen	81098722d2	add chinese version of humaneval, mbpp (#743 ) * add chinese_version of humaneval,mbpp * add humaneval&mbpp gen.py * minor fix * minor add --------- Co-authored-by: yingfhu <yingfhu@gmail.com>	2023-12-28 14:47:56 +08:00
bittersweet1999	db919f0191	[Fix] SubSizePartition fix (#746 ) * fix subjective_eval * subject_eval partition situation fixed * subject_eval partition situation fixed	2023-12-28 11:46:46 +08:00
Hubert	0a525985e8	[Feature] Support sanitized MBPP dataset (#745 )	2023-12-27 22:17:23 +08:00
bittersweet1999	dfd9ac0fd9	[Feature] Add other judgelm prompts for Alignbench (#731 ) * add judgellm prompts * add judgelm prompts * update import info * fix situation that no abbr in config * fix situation that no abbr in config * add summarizer for other judgellm * change config name * add maxlen * add maxlen * dict assert * dict assert * fix strings * fix strings	2023-12-27 17:54:53 +08:00
Yang Yong	54345c56b7	Update LightllmApi and Fix mmlu bug (#738 ) * Update LightllmApi and Fix mmlu bug * checkout mmlu_gen_a484b3.py --------- Co-authored-by: Leymore <zfz-960727@163.com>	2023-12-27 13:49:08 +08:00
philipwangOvO	34561ececb	[Feature] Add InfiniteBench (#739 ) * add InfiniteBench * add InfiniteBench --------- Co-authored-by: wangchonghua <wangchonghua@pjlab.org.cn>	2023-12-26 15:36:27 +08:00
Fengzhe Zhou	3a68083ecc	[Sync] update configs (#734 )	2023-12-25 21:59:16 +08:00
AllentDan	336d8d76ff	add turbomind restful api support (#693 ) * add turbomind restful api support * config * top_p 0.8 * top_k = 1	2023-12-24 01:40:00 +08:00
bittersweet1999	e985100cd1	[Fix] Fix subjective alignbench (#730 )	2023-12-23 20:06:53 +08:00
Mo Li	0e24f4213e	[Feature] Add NeedleInAHaystack Test Support (#714 ) * Add NeedleInAHaystack Test * Apply pre-commit formatting * Update configs/eval_hf_internlm_chat_20b_cdme.py Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com> * add needle in haystack test * update needle in haystack test --------- Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com>	2023-12-23 12:00:51 +08:00
RunningLeon	e34c552282	[Feature] Update configs for evaluating chat models like qwen, baichuan, llama2 using turbomind backend (#721 ) * add llama2 test * fix * test qwen chat-7b * test w4 * add baichuan2 * update * update * update configs and docs * update	2023-12-21 18:22:17 +08:00
bittersweet1999	fbb912ddf3	[Feature] Add abbr for judgemodel in subjective evaluation (#724 ) * add_judgemodel_abbr * add judgemodel abbr	2023-12-21 15:58:20 +08:00
Skyfall-xzz	b35d991786	[Feature] Add ReasonBench(Internal) dataset (#577 ) * [Feature] Add reasonbench dataset * add configs for supporting generative inference & merge datasets in the same category * modify config filename to prompt version * fix codes to meet pre-commit requirements * lint the code to meet pre-commit requirements * Align Load_data Sourcecode Briefly * fix bugs * reduce code redundancy	2023-12-20 17:57:42 +08:00
Jingming	76a95e9e81	[Feature] Support the use of humaneval_plus. (#720 ) * [Feature] Support the use of humaneval_plus. * [Feature] Add humaneval_plus_gen.py * minor check * [Fix] Fix bug --------- Co-authored-by: yingfhu <yingfhu@gmail.com>	2023-12-20 17:25:17 +08:00
bittersweet1999	97c2068bd9	[Feature] Add JudgeLLMs (#710 ) * add judgellms * add judgellms * add sub_size_partition * add docs * add ref	2023-12-19 18:40:25 +08:00
Hubert	eda72e756e	[Fix] minor fix openai (#711 )	2023-12-18 15:45:31 +08:00
Songyang Zhang	637628a70f	[Doc] Update Doc for Alignbench (#707 ) * update alignmentbench * update alignmentbench * update doc * update * update	2023-12-15 15:07:25 +08:00
DseidLi	db2920326a	[Fix] remove redundant in gsm8k.py (#700 ) Removed redundant code in GSM8KDataset.load method.	2023-12-14 19:55:58 +08:00
Songyang Zhang	bfe4aa2af5	[Fix] Update alignmentbench (#704 ) * update alignmentbench * update alignmentbench * update alignmentbench	2023-12-14 18:24:21 +08:00
bittersweet1999	1fe152b3e8	[Feature] Support AlignmentBench infer and judge (#697 ) * alignmentbench infer and judge * alignmentbench * alignmentbench done * alignment all done * alignment all done	2023-12-13 19:59:30 +08:00
Hubert	a94598d921	[Feat] update python action and slurm (#694 )	2023-12-13 10:41:10 +08:00
bittersweet1999	6130394165	[Feature] Add double order of subjective evaluation and removing duplicated response among two models (#692 ) * add features * add doc string * add doc string	2023-12-12 20:58:17 +08:00
Hubert	4780b39eda	[Sync] format (#690 ) Co-authored-by: Leymore <zfz-960727@163.com>	2023-12-12 14:03:45 +08:00
bittersweet1999	3e77175720	[Fix] Hotfix for Subjective Evaluation (#686 )	2023-12-12 09:22:08 +08:00
bittersweet1999	465308e430	[Feature] Add Subjective Evaluation (#680 ) * new version of subject * fixed draw * fixed draw * fixed draw * done * done * done * done * fixed lint	2023-12-11 22:22:11 +08:00
Hubert	4f0b373a0a	[Fix] fix docstring (#684 )	2023-12-11 19:12:01 +08:00
Hubert	e78857ac36	[Sync] minor test (#683 )	2023-12-11 17:42:53 +08:00
Jingming	dd4318f6ab	[Feature] enhance the ability of humaneval_postprocess (#676 ) * [Feature] enhance the ability of humaneval_postprocess * refactor * [Feature] Keep the old version of the function and realize the new function in humaneval_postprocess_v2. * Update opencompass/datasets/humaneval.py --------- Co-authored-by: Leymore <zfz-960727@163.com> Co-authored-by: Hubert <42952108+yingfhu@users.noreply.github.com>	2023-12-11 14:39:56 +08:00
Songyang Zhang	e25c5f9525	[Enhancement] Update API Interface and Mixtral (#681 ) * [Enhancement] Update API interface * [Enhancement] Update API interface * Update mixtral * Update readme	2023-12-10 13:29:26 +08:00
Xiaoming Shi	1bf85949ef	[Feature] Add medbench (#678 ) * update medbench * medbench update * format medbench * format --------- Co-authored-by: 施晓明 <PJLAB\shixiaoming@pjnl104220118l.pjlab.org> Co-authored-by: Leymore <zfz-960727@163.com>	2023-12-09 16:05:46 +08:00
Jingming	7cb53a95fa	[Fix] fix bug on standart_deviation summarizer (#675 )	2023-12-08 13:38:07 +08:00
liyucheng09	05bbce8b08	[Feature] Add Data Contamination Analysis (#639 ) * add contamination analysis to ceval * fix bugs * add contamination docs * to pass CI check * update --------- Co-authored-by: zhangyifan1 <zhangyifan1@pjlab.org.cn> Co-authored-by: Leymore <zfz-960727@163.com>	2023-12-08 10:00:11 +08:00
bittersweet1999	1c95790fdd	New subjective judgement (#660 ) * TabMWP * TabMWP * fixed * fixed * fixed * done * done * done * add new subjective judgement * add new subjective judgement * add new subjective judgement * add new subjective judgement * add new subjective judgement * modified to a more general way * modified to a more general way * final * final * add summarizer * add new summarize * fixed * fixed * fixed --------- Co-authored-by: caomaosong <caomaosong@pjlab.org.cn>	2023-12-06 13:28:33 +08:00
rolellm	e10f1c9139	added rolebench dataset. (#633 ) * added rolebench * 修改了不合理的变量名 * 修改了评论中的变量名	2023-12-01 22:54:42 +08:00
Hubert	9eb5cadcac	[Feat] update gsm8k and math agent config (#652 ) * [Feat] update gsm8k and math agent config * minor fix	2023-12-01 15:08:38 +08:00
liushz	a331c9abfd	[Feature] Add wikibench dataset (#655 ) * Add WikiBench * Add WikiBench * format --------- Co-authored-by: Leymore <zfz-960727@163.com>	2023-12-01 14:56:54 +08:00
liushz	e019c831fe	[Feature] Add Chinese version: commonsenseqa, crowspairs and nq (#144 ) * add Chinese version: csqa crowspairs nq * Update cn_data * Update cn_data * update format --------- Co-authored-by: liuhongwei <liuhongwei@pjlab.org.cn> Co-authored-by: Leymore <zfz-960727@163.com>	2023-11-30 15:33:02 +08:00
Ma Zerun	6aaf3b91ec	[Feature] Support chat style inferencer. (#643 ) * [Feature] Support chat style inferencer. * [Fix] use new prompt * [Fix] use new prompt --------- Co-authored-by: yingfhu <yingfhu@gmail.com>	2023-11-30 14:00:06 +08:00
Fengzhe Zhou	e20d654c18	[Sync] Bump version to 0.1.9 (#644 )	2023-11-28 11:42:43 +08:00
Hubert	d4af31bab4	[Feat] support zhipu post process (#642 ) * [Feat] support zhipu post * [Feat] support zhipu post * [Feat] support zhipu post	2023-11-27 19:57:36 +08:00
liushz	6d0d78986c	[Feature] Add GSM_Hard dataset (#619 ) * Add SVAMP dataset * Add SVAMP dataset * Add SVAMP dataset * Add gsm_hard dataset * Add gsm_hard dataset * format --------- Co-authored-by: Leymore <zfz-960727@163.com>	2023-11-27 17:40:34 +08:00
Fengzhe Zhou	9083dea683	[Sync] some renaming (#641 )	2023-11-27 16:06:49 +08:00
Yang Yong	522241a8c9	[Fix] Fix lightllmapi list bug (#635 )	2023-11-24 14:24:13 +08:00
Hubert	1884912674	[Bug] fix icl eval with nested list (#632 )	2023-11-24 13:43:26 +08:00
Fengzhe Zhou	79f6449d85	[Doc] Update FAQ (#628 ) * update faq * Update docs/zh_cn/get_started/faq.md * Update docs/en/get_started/faq.md * Update docs/zh_cn/get_started/faq.md --------- Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com>	2023-11-23 18:19:17 +08:00
Fengzhe Zhou	d949e3c003	[Feature] Add circular eval (#610 ) * refactor default, add circular summarizer * add circular * update impl * update doc * minor update * no more to be added	2023-11-23 16:45:47 +08:00
Songyang Zhang	5202456b4c	[API] Update API (#624 ) * update api * update generation_kwargs impl * update api * refactor --------- Co-authored-by: Leymore <zfz-960727@163.com>	2023-11-23 15:06:20 +08:00
Fengzhe Zhou	d4d1330a5a	[Sync] Fix cmnli, fix vicuna meta template, fix longbench postprocess and other minor fixes (#625 )	2023-11-23 14:05:59 +08:00
Kevin Wang	c0785e53d8	[Feature] support download from modelscope (#534 ) * [Feature] download from modelscope * [Feature] download from modelscope * minor fix --------- Co-authored-by: yingfhu <yingfhu@gmail.com>	2023-11-22 15:32:21 +08:00
liushz	048775192b	[Feature] Add SVAMP dataset (#604 ) * Add SVAMP dataset * Add SVAMP dataset * Add SVAMP dataset	2023-11-22 14:54:39 +08:00
Fengzhe Zhou	fb30b7c7a2	[Fix] Fix gen inferencer (#615 )	2023-11-22 12:04:31 +08:00
Songyang Zhang	721a45c68f	[Bug] Update api with generation_kargs (#614 ) * update api * update generation_kwargs impl --------- Co-authored-by: Leymore <zfz-960727@163.com>	2023-11-22 10:02:57 +08:00
Lyu Han	eb56fd6d16	Integrate turbomind python api (#484 ) * integrate turbomind python api * update * update user guide * update * fix according to reviewer's comments * fix error * fix linting * update user guide * remove debug log --------- Co-authored-by: Songyang Zhang <tonysy@users.noreply.github.com>	2023-11-21 22:34:46 +08:00
Songyang Zhang	d925748266	[Feature] Support 360API and FixKRetriever for CSQA dataset (#601 ) * [Feature] Support 360API and FixKRetriever for CSQA dataset * Update API * Update API * [Feature] Support 360API and FixKRetriever for CSQA dataset * Update API * Update API * rm mathbench * fix_lint * Update opencompass/models/bytedance_api.py Co-authored-by: Hubert <42952108+yingfhu@users.noreply.github.com> * update * update * update --------- Co-authored-by: Hubert <42952108+yingfhu@users.noreply.github.com>	2023-11-21 20:25:47 +08:00
Yang Yong	d3b0d5c4ce	[Feature] Support Lightllm API (#613 ) * [Feature] Support Lightllm api * formatting & renaming --------- Co-authored-by: Leymore <zfz-960727@163.com>	2023-11-21 19:18:40 +08:00
Yuan Feng	7199acc25d	Add support for DataCanvas Alaya LM (#612 ) * Support for Alaya * Remove useless requirements	2023-11-21 17:51:30 +08:00
liushz	dbacd36379	Add aritch to mathbench (#607 )	2023-11-20 19:40:41 +08:00
liushz	c9c5c5d92e	Mathbench update postprocess (#600 ) * Update mathbench * Update mathbench	2023-11-20 16:48:55 +08:00
Jingming	5e75e29711	[Feature] Add multi-prompt generation demo (#568 ) * [Feature] Add multi-prompt generation demo * [Fix] change form in winogrande_gen_XXX.py * [Fix] make multi prompt demo more directly * [Fix] fix bug * [Fix] minor fix --------- Co-authored-by: yingfhu <yingfhu@gmail.com>	2023-11-20 16:16:37 +08:00
Hubert	91fba2c2e9	[Feat] support humaneval and mbpp pass@k (#598 ) * [Feat] support pass@ k * [Feat] support pass@k * [Feat] support pass@k * [Feat] support pass@k * [Feat] support pass@k * [Feat] support pass@k docs * update naming --------- Co-authored-by: Leymore <zfz-960727@163.com>	2023-11-16 21:22:06 +08:00
Raymond Zhang	c0acd06b05	[Feature] Add FinanceIQ dataset (#596 )	2023-11-16 17:47:57 +08:00
Hubert	fcab30f82e	[Fix] change save_every defaults to 1 (#592 )	2023-11-15 13:00:25 +08:00
Fengzhe Zhou	19ad7f9613	fix cmb dataset (#587 )	2023-11-14 16:13:39 +08:00
Wei Jueqi	14e6fe6f13	Fix bugs in subjective evaluation (#589 ) * rename * fix sub bugs and update docs * update * update	2023-11-14 16:11:55 +08:00
Fengzhe Zhou	1ea88d5822	[Sync] Bump version to 0.1.8 (#576 )	2023-11-13 16:00:38 +08:00
Fengzhe Zhou	d3de5c41fb	[Sync] update model configs (#574 )	2023-11-13 15:15:34 +08:00
Fengzhe Zhou	689ffe5b63	[Feature] Use dataset in local path (#570 ) * update commonsenseqa * update drop * update flores_first100 * update gsm8k * update humaneval * update lambda * update obqa * update piqa * update race * update siqa * update story_cloze * update strategyqa * update tydiqa * update winogrande * update doc * update hellaswag * fix obqa * update collections * update .zip name	2023-11-13 13:00:37 +08:00
Fengzhe Zhou	d6aaac22e7	[Feature] Update cmb (#571 )	2023-11-13 00:09:05 +08:00
Songyang Zhang	9e42cb163b	[Feature] Update xunfei api (#572 ) * update xunfei api * fix lint * avoid warning	2023-11-10 22:46:06 +08:00
jingmingzhuo	b3cbef3226	[Feature] Add py150 and maxmin (#562 ) * [feat] add clozeTesst_maxmin dataset * [feat] add py150 datasets * [feat] change __init__.py in opencompass/datasets * [fix] pre-commit check * [fix] rename py150 and masxmin datasets in configs * [feat] add gen.py of py150 and maxmin in configs/datasets	2023-11-09 22:05:25 +08:00
Hubert	889a6b26ae	[Fix] fix log re-direct (#564 )	2023-11-09 19:34:19 +08:00
Hubert	cf5a6d1ab7	[Fix] fix unnecessary import and update requirements (#555 )	2023-11-08 17:58:49 +08:00
Hubert	9f8a721313	[Fix] fix registry error with internal (#551 ) * [Fix] fix conflict with internal * [Fix] fix conflict with internal	2023-11-07 20:01:23 +08:00
Hubert	bb2ecf416e	[Feat] Support cibench (#538 ) * [Feat] support cidataset * [Feat] support cidataset * [Feat] support cidataset * [Feat] support cidataset * minor fix * minor fix * minor fix * minor fix * minor fix * minor fix * rename cibench * rename cibench * rename cibench * rename cibench * minor fix * minor fix * minor fix	2023-11-07 19:11:44 +08:00
Songyang Zhang	239c2a346e	[Feature] Add support for MiniMax API (#548 ) * update requirement * update requirement * update with minimax * update api model * Update readme * fix error --------- Co-authored-by: zhangsongyang <zhangsongyang@pjlab.org.cn>	2023-11-06 21:57:32 +08:00
Hubert	1ccdfaa623	[Feat] support xunfei api (#547 )	2023-11-06 19:29:26 +08:00
Yuan Liu	6e31520128	[Feature]: To be compatible with the latest version of MiniGPT-4 (#539 ) * [Feature]: To be compatible with the latest version of MiniGPT-4 * [Feature]: User try and except Co-authored-by: Fengzhe Zhou <zfz-960727@163.com> * [Fix]: Fix lint --------- Co-authored-by: bensenliu <bensenliu@tencent.com> Co-authored-by: Fengzhe Zhou <zfz-960727@163.com>	2023-11-04 09:50:36 +08:00
bittersweet1999	f25a980043	[fFeat] Add an opensource dataset Tabmwp (#505 ) * TabMWP * TabMWP * fixed * fixed * fixed * done * done * done --------- Co-authored-by: caomaosong <caomaosong@pjlab.org.cn>	2023-11-03 11:15:46 +08:00
Hubert	b9270c3a60	[Fix] Fix local debug mode not restrict the resources (#522 ) * [Fix] fix local debug mode not restrict the resources * minor fix	2023-10-30 18:13:43 +08:00
Qing	e2355a2ede	[Feature] Add multi model viz (#509 ) * add viz_multi_model.py tool * Modify the viz_multi_model.py script according to the review * highlight multiple optimal scores --------- Co-authored-by: wq.chu <wq.chu@tianrang-inc.com> Co-authored-by: Leymore <zfz-960727@163.com>	2023-10-30 12:11:33 +08:00
Fengzhe Zhou	6a398d171c	Bump version to 0.1.7 (#518 )	2023-10-27 20:32:27 +08:00
Fengzhe Zhou	dbb20b8270	[Sync] update (#517 )	2023-10-27 20:31:22 +08:00
Hubert	6f07af3039	[Feat] Support local runner for windows (#515 )	2023-10-27 17:16:22 +08:00
Fengzhe Zhou	df07391ed8	[Fix] Enforce `do_sample=False` in HF model (#506 ) * update hf model wrapper * patch llama --------- Co-authored-by: bot <bot@bot.com>	2023-10-27 16:54:19 +08:00
Wei Jueqi	b62842335d	[Doc] Update Subjective docs (#510 ) * rename * add en subdoc * fix name * fix writing * update --------- Co-authored-by: Leymore <zfz-960727@163.com>	2023-10-27 16:27:24 +08:00
Fengzhe Zhou	e3d4901bed	[Feat] Add _set_model_kwargs_torch_dtype for HF model (#507 ) * add _set_model_kwargs_torch_dtype for hf models * add logger	2023-10-27 11:45:41 +08:00
Fengzhe Zhou	6405cd2db5	use example summarizer by default (#508 )	2023-10-27 11:45:29 +08:00
Hubert	b3f5d9e421	[Feat] support math/gms8k agent config (#494 ) * support math agent * support gsm8k agent * support gsm8k agent * minor fix * minor fix * minor fix * Update configs/eval_codeagent.py	2023-10-25 23:05:15 +08:00
Hubert	ac3a2c4501	[Feat] local api speed up with fixed concurrent users (#497 ) * [Feat] local api speed up * fix lint * fix lint * minor fix * add example api	2023-10-25 21:12:20 +08:00
Leymore	4dd9a3fc10	[Sync] sync with internal codes 20231019 (#488 )	2023-10-18 23:37:35 -05:00
liushz	2737249f31	[Feature] Add mathbench dataset and circular evaluator (#408 ) * add_mathbench * update mathbench * support non circular eval dataset --------- Co-authored-by: liuhongwei <liuhongwei@pjlab.org.cn> Co-authored-by: yingfhu <yingfhu@gmail.com>	2023-10-18 04:08:31 -05:00
Leymore	fccfcb6f5b	fix summary default (#483 )	2023-10-17 11:32:38 +08:00
Leymore	6317da08b3	Bump version to 0.1.6 (#478 )	2023-10-13 06:54:51 -05:00
Leymore	7d9e386821	[Fix] Split if and only if complete eos string shows up (#477 )	2023-10-13 06:52:20 -05:00
Leymore	861942ab1b	[Feature] Add lawbench (#460 ) * add lawbench * update requirements * update	2023-10-13 06:51:36 -05:00
Leymore	fbf5089c40	[Sync] update github token (#475 )	2023-10-13 06:50:54 -05:00
Leymore	362c33dff4	fix jieba rouge (#467 )	2023-10-12 10:25:19 +08:00
Leymore	d7ff933a73	[Fix] Use jieba rouge in lcsts (#459 ) * use jieba rouge in lcsts * use rouge_chinese	2023-10-09 10:10:33 +08:00
Tong Gao	119bfd1569	[Refactor] Move fix_id_list to Retriever (#442 ) * [Refactor] Move fix_id_list to Retriever * update * move to base * fix	2023-10-07 12:53:41 +08:00
Lyu Han	6738247142	Integrate turbomind inference via its RPC API instead of its python API (#414 ) * support tis * integrate turbomind inference via its RPC API instead of its python API * update guide * update ip address spec * update according to reviewer's comments	2023-10-07 10:27:48 +08:00
Leymore	9db5652638	[Feature] re-implement ceval load dataset (#446 )	2023-09-27 21:18:48 +08:00
Hubert	d9f3e88dfe	[Fix] fix clp potential error and support bs>1 (#439 ) * [Fix] fix clp potential error and support bs>1 * [Fix] fix clp potential error and support bs>1 * minor fix * minor fix	2023-09-27 16:32:57 +08:00
philipwangOvO	3bb3d330eb	[Sync] Update LongEval (#443 )	2023-09-27 16:32:40 +08:00
Tong Gao	9b21613d17	Bump version to 0.1.5 (#432 )	2023-09-22 19:17:23 +08:00
chenbohua3	b2926eac8f	[Feature] support customize config path (#423 ) * support customize config path * support customize config path * support customize config path	2023-09-22 19:12:02 +08:00
liushz	c5224c2a91	[Feature] Add kaoshi dataset (#392 ) * Add ToT method * Update ToT * Update ToT * Update ToT * Update ToT * Update ToT * Add Koashi * Update Kaoshi * Update Kaoshi * Update kaoshi * Update kaoshi * Update Kaoshi * Update Kaoshi * Update Kaoshi * Update Kaoshi * update Kaoshi * update * update * fix --------- Co-authored-by: gaotongxiao <gaotongxiao@gmail.com>	2023-09-22 18:46:33 +08:00
TTTTTiam	2a62bea1a4	add evaluation of scibench (#393 ) * add evaluation of scibench * add evaluation of scibench * update scibench * remove scibench evaluator --------- Co-authored-by: Leymore <zfz-960727@163.com>	2023-09-22 17:42:08 +08:00
Tong Gao	07574fddbb	[Fix] keep keys (#431 )	2023-09-22 17:30:54 +08:00
Tong Gao	a1ea3c094a	[Sync] Initial support of subjective evaluation (#421 ) Co-authored-by: Leymore <zfz-960727@163.com>	2023-09-22 15:42:31 +08:00
Ma Zerun	0f2c388280	Support GSM8k evaluation with tools by Lagent and LangChain (#277 ) * Support GSM8k evaluation with tools by Lagent and LangChain * Avoid to use MMEngine new feature * update document --------- Co-authored-by: Leymore <zfz-960727@163.com>	2023-09-22 15:28:22 +08:00
Tong Gao	681d3013de	[Feature] Log gold answer in prediction output (#419 ) * [Feature] Log gold answer in prediction output * support clp golden ans * minor fix --------- Co-authored-by: yingfhu <yingfhu@gmail.com>	2023-09-22 12:44:40 +08:00
Yike Yuan	97fdc51102	[Fix] Fix performance issue of visualglm. (#424 ) * [Fix] Visualglm performance fixed. * [Fix] Hide ckpt path.	2023-09-21 19:54:23 +08:00
Hubert	8803f7f7a6	[Feat] support antropics evals dataset (#422 ) * [Feat] support anthropics ai risk dataset * [Feat] support anthropics evals dataset * [Feat] support anthropics evals dataset	2023-09-20 18:36:44 +08:00
Leymore	ae0cd8752f	[Feature] Use local accuracy from hf implements (#416 ) * use local accuracy from hf implements * add load from hf fallback	2023-09-20 16:35:22 +08:00
Zequn Liu	ff2c15a09f	[fix] summarizer debug logger (#417 )	2023-09-20 15:29:26 +08:00
Yike Yuan	bd50bad8b5	[Feat] Support mm models on public dataset and fix several issues. (#412 ) * [Feat] Add public dataset support for visualglm, qwenvl, and flamingo * [Fix] MMBench related changes. * [Fix] Openflamingo inference. * [Fix] Hide ckpt path. * [Fix] Pre-commit. --------- Co-authored-by: Haodong Duan <dhd.efz@gmail.com>	2023-09-19 19:08:44 +08:00
Yuanhan Zhang	7c2726c23b	[Model] Yhzhang/add mlugowl llamaadapter (#405 ) * refine gitignore * [Feature]: Add minigpt-4 * [Feature]: Add mm local runner * [Feature]: Add instructblip * add otter and llama-adapter * add owl * add llama2-adapter and owl * lint * [Feature]: Add minigpt-4 * [Feature]: Add instructblip * add otter and llama-adapter * add owl * add llama2-adapter and owl * lint * lint * update * lint * lint * add __init__.py * update * update * update * update * [Feature]: Add minigpt-4 * [Feature]: Add mm local runner * [Feature]: Add instructblip * add otter and llama-adapter * add owl * add llama2-adapter and owl * lint * [Feature]: Add minigpt-4 * [Feature]: Add instructblip * add otter and llama-adapter * add owl * add llama2-adapter and owl * lint * lint * update * lint * lint * add __init__.py * update * update * update * update * optimize mmbench dataset args * update * update * run commit hook --------- Co-authored-by: liuyuan <3463423099@qq.com> Co-authored-by: kennymckormick <dhd@pku.edu.cn> Co-authored-by: kennymckormick <dhd.efz@gmail.com>	2023-09-19 14:21:26 +08:00
so2liu	267401bded	[Feat] add custom summarizer argument in CLI run mode 在CLI启动模式中添加自定义Summarizer参数 (#411 ) * feat: add custom summarizer in CLI run mode * feat: search local config by match_cfg_file	2023-09-18 18:11:22 +08:00
Hubert	2c15a0c01d	[Feat] refine docs and codes for more user guides (#409 )	2023-09-18 16:12:13 +08:00
Hubert	a11cb45c83	[Feat] implementation for support promptbench (#239 ) * [Feat] support adv_glue dataset for adversarial robustness * reorg files * minor fix * minor fix * support prompt bench demo * minor fix * minor fix * minor fix * minor fix * minor fix * minor fix * minor fix * minor fix	2023-09-15 15:06:53 +08:00
Hubert	de8a154795	[Feat] support ds1000 dataset (#395 ) * [Feat] support ds1000 datase	2023-09-15 12:50:27 +08:00
Xidong Wang	47a752cd56	[Dataset] Add CMB (#376 ) * Add CMB * modify CMB --------- Co-authored-by: wangxidong <xidongw@163.com>	2023-09-12 19:16:41 +08:00
cdpath	722eb39526	fix potential oom issue (#387 )	2023-09-12 10:41:03 +08:00
Tong Gao	c7a8b8fe98	Bump version to 0.1.4 (#367 )	2023-09-08 20:51:38 +08:00
Yixiao Fang	fada77a31c	[Feature] Add open source dataset eval config of instruct-blip (#370 ) * add configs * refactor model * add post processor and prompt constructor	2023-09-08 15:07:09 +08:00
Leymore	49c467458f	[Feature] Update llama2 (#372 )	2023-09-08 12:47:56 +08:00
Tong Gao	b11838f80a	[Feature] Update claude2 postprocessor (#365 ) * [Feature] Update claude2 config * [Feature] Update claude2 postprocessor	2023-09-07 11:26:26 +08:00
Yike Yuan	b885ec84df	[Feat] Support Qwen-VL-Chat on MMBench. (#312 ) * [Feat] Support Qwen-VL base. * [Feat] Support Qwen-VL-Chat on MMBench. * [Fix] Add postprocessor and fix format. * [Fix] Add type hint and remove redundant codes. * [Fix] fix bugs in postprocessor. * [Fix] Use given commit id.	2023-09-06 18:42:19 +08:00
Hubert	ddb8197212	[Feat] support wizardcoder series (#344 ) * [Feat] support wizardcoder series * minor fix	2023-09-06 17:52:35 +08:00
Leymore	880b34e759	[Fix] Quick lint fix (#362 ) * add default value * lint fix * use None	2023-09-06 14:33:13 +08:00
Tong Gao	5d75c1bbb9	[Enhancement] Increase default task size (#360 )	2023-09-05 10:38:13 +08:00
Leymore	b8bf16e81c	[Fix] zero retriever add default value (#361 )	2023-09-05 10:37:42 +08:00
Mashiro	ab21f3be66	[Enhance] Supress warning raised by get_logger (#353 )	2023-09-04 15:27:08 +08:00
Leymore	a1782f9a08	[Fix] triviaqa & nq postprocess (#350 )	2023-09-04 15:24:52 +08:00
Tong Gao	ce65d3393b	[Sync] Use finally to clean up temp files (#337 )	2023-09-04 15:20:16 +08:00
Yixiao Fang	2cd994c3d1	[Fix] add import check of multimodal (#352 )	2023-09-04 14:41:07 +08:00
Leymore	8774465a8f	[Enhancement] ignore ZeroRetriever error when id_list provided (#340 )	2023-09-04 11:12:16 +08:00
Yuanhan Zhang	f2dd98ca7a	[Feat] Support LLaVA and mPLUG-Owl (#331 ) * refine gitignore * [Feature]: Add minigpt-4 * [Feature]: Add mm local runner * [Feature]: Add instructblip * add otter and llama-adapter * add owl * add llama2-adapter and owl * lint * [Feature]: Add minigpt-4 * [Feature]: Add instructblip * add otter and llama-adapter * add owl * add llama2-adapter and owl * lint * lint * update * lint * lint * add __init__.py * update * update * update --------- Co-authored-by: liuyuan <3463423099@qq.com>	2023-09-01 23:32:05 +08:00
Leymore	e810974068	[Fix] Fix when missing both pad and eos token (#287 ) * fix when missing both pad and eos token * update pad_token_id impl	2023-08-31 16:53:39 +08:00
Li Bo	a4d6840739	[Feat] Add Otter to OpenCompass MMBench Evaluation (#232 ) * add otter model for opencompass mmbench * add docs * add readme docs * debug for otter opencomass eval * delete unused folders * change to default data path * remove unused files * remove unused files * update * update config file * flake8 lint formated and add prompt generator * add prompt generator to config * add a specific postproecss * add post processor * add post processor * add post processor * update according to suggestions * remove unused redefinition	2023-08-31 12:55:53 +08:00
Leymore	7ca6ba625e	[Feature] Add qwen & qwen-chat support (#286 ) * add and apply update suffix tool * add tool doc * add qwen configs * add cmmlu * rename bbh * update datasets * delete * update hf_qwen_7b.py	2023-08-31 11:29:05 +08:00
Hubert	fd389e2d78	[Feat] support codellama and preds collection tools (#335 )	2023-08-31 11:14:42 +08:00
Tong Gao	9058be07b8	[Feature] Simplify entry script (#204 ) * [Feature] Simply entry script * update	2023-08-25 17:36:30 +08:00
Tong Gao	f480b72703	[Feature] Support model-bound prediction postprocessor, use it in Claude (#268 ) * [Feature] Support model-bound text postprocessor, add claude as an example * update * update * minor fix --------- Co-authored-by: zhoufengzhe <zhoufengzhe@pjlab.org.cn>	2023-08-25 16:12:21 +08:00
Yike Yuan	3f601f420b	[Feat] Support public dataset of visualglm and llava. (#265 ) * [Feat] Add public dataset support of VisualGLM. * [Feat] Refactor LLaVA. * [Feat] Add public dataset support of LlaVA. * [Fix] Add arg.	2023-08-25 15:44:32 +08:00
Yuan Liu	dc6e54f6f4	[Feature]: Verify the acc of these public datasets (#269 ) * [Feature]: Refactor public dataset eval * [Feature]: Verify public dataset acc	2023-08-25 15:01:58 +08:00
philipwangOvO	3f37c40aa3	[Dataset] Refactor LEval	2023-08-25 11:46:23 +08:00
Tong Gao	60c2d3d76b	[Feature] Add Claude support (#253 ) * [Feature] Add Claude support * [Feature] Add Claude support * Update opencompass/models/claude_api.py Co-authored-by: Hubert <42952108+yingfhu@users.noreply.github.com> * raise import erorr --------- Co-authored-by: Hubert <42952108+yingfhu@users.noreply.github.com>	2023-08-24 14:29:45 +08:00
Yuan Liu	343f785b07	[Feature]: Add Flamingo (#258 ) * [Feature]: Add Openflamingo MMBench * [Fix]: Fix import error * [Fix]: Revert task config * [Fix]: Fix path bug	2023-08-24 14:11:29 +08:00
LZHgrla	77745a84ea	[Fix] Fix bugs for PeftModel generate (#252 ) * fix bugs * fix typo	2023-08-24 14:07:33 +08:00
Tong Gao	bd47a00f27	[Fix] use sympy only when necessary (#255 )	2023-08-24 10:15:20 +08:00
Tong Gao	01372a4806	update (#251 )	2023-08-23 16:25:23 +08:00
Yixiao Fang	1034c487ef	[Refactor] Refactor instructblip (#227 ) * refactor instructblip * add post processor * add forward * fix lint * update * update	2023-08-23 15:33:59 +08:00
liushz	02ce139bc6	[Feature] Add Tree-of-Thought method (#173 ) * Add ToT method * Update ToT * Update ToT * Update ToT * Update ToT * Update ToT * Update ToT * Update ToT * Update chain_of_thought.md * Update icl_tot_inferencer.py --------- Co-authored-by: liuhongwei <liuhongwei@pjlab.org.cn>	2023-08-23 12:23:05 +08:00
Leymore	ff5ab92331	[Feature] Add llama2 native implements (#235 ) * add llama2 native implements * rename configs/eval_llama_7b.py --------- Co-authored-by: zhoufengzhe <zhoufengzhe@pjlab.org.cn>	2023-08-23 11:33:25 +08:00
Leymore	fdc69f9d58	[Fix] local runner debug (#238 )	2023-08-21 16:58:36 +08:00
Yike Yuan	8d368d1cd6	[Feat] Support visualglm and llava for MMBench evaluation. (#211 ) * [Feat] Support visualglm inference on MMBench. * [Feat] Support llava inference on MMBench. * [Fix] Fix pre-commit format. * [Fix] Add docstring for llava * [Fix] Fix multi-process inference error of LlaVA and add comments. 1. Set `low_cpu_mem_usage` to False to address device issue. 2. Add docstring and type hints. 3. Rename class and remove registry. * [Fix] Pre-commit fix. * [Fix] add forward entry, add dynamic import to seedbench * [Fix] Fix pre-commit. * [Fix] Fix missing context. * [Fix] Fix docstring.	2023-08-21 15:57:30 +08:00
Yike Yuan	a6552224cb	[Feat] Support multi-modal evaluation on MME benchmark. (#197 ) * [Feat] Support multi-modal evaluation on MME benchmark. * [Fix] Remove debug code. * [Fix] Remove redundant codes and add type hints. * [Fix] Rename in config. * [Fix] Rebase main. * [Fix] Fix isort and yapf conflict.	2023-08-21 15:53:20 +08:00
philipwangOvO	3b29aaee2b	[Fix] bin_trim (#237 ) Co-authored-by: wangchonghua <wangchonghua@pjlab.org.cn>	2023-08-21 15:44:49 +08:00
philipwangOvO	655a807f4b	[Dataset] LongBench (#236 ) Co-authored-by: wangchonghua <wangchonghua@pjlab.org.cn>	2023-08-21 14:15:20 +08:00
Yixiao Fang	0fa2482661	[Feature] Support SEED-Bench (#203 ) * support seedbench * update docstrings * update * update * update * update according to review * rebase * fix lint * update	2023-08-17 17:24:02 +08:00
Ezra-Yu	17ccaa5980	[Feat] Add codegeex2 and Humanevalx (#210 ) * add codegeex2 * add humanevalx dataset * add evaluator * update evaluator * update configs * update clean code * update configs * fix lint * remove sleep * fix lint * update docs * fix lint	2023-08-17 11:03:16 +08:00
Hubert	0fe2366a72	[Feat] support adv_glue dataset for adversarial robustness (#205 ) * [Feat] support adv_glue dataset for adversarial robustness * reorg files * minor fix * minor fix	2023-08-16 18:42:06 +08:00
Yuan Liu	78df9bd0cb	[Feature]: Add other public datasets (#206 ) * [Feature]: Refactor class name * [Feature]: Add minigpt-4 coco caption * [Feature]: Update minigpt-4 coco caption * [Feature]: Add MiniGPT-4 ScienceQA * [Feature]: Add minigpt-4 vqav2 * [Feature]: Add VSR * [Feature]: Revert task to previous version	2023-08-16 11:37:26 +08:00
Yike Yuan	3a46b6c64f	[Fix] Fix bugs of multiple rounds of inference when using mm_eval (#201 )	2023-08-16 11:15:11 +08:00
Hubert	7c393192af	[Fix] fix bug for postprocessor (#195 ) * [Fix] fix bug for postprocessor * minor fix	2023-08-11 18:41:12 +08:00
Tong Gao	10cbc2b175	Bump version to 0.1.2 (#190 )	2023-08-11 17:43:14 +08:00
Tong Gao	bf79ff1c6d	[Feature] Add LEval datasets Co-authored-by: kennymckormick <dhd@pku.edu.cn>	2023-08-11 17:38:31 +08:00
Hubert	8d9cee060f	[Feat] update postprocessor to get first option more accurately (#193 ) * [Feat] update postprocessor to get first option * minor fix * minor fix	2023-08-11 17:33:00 +08:00
Leymore	14332e08fd	[Feature] add llama-oriented dataset configs (#82 ) * add llama-oriented dataset configs * update * revert cvalues & update llama_example	2023-08-11 12:48:05 +08:00
Hubert	5a9539f375	[Feat] add safety to collections (#185 ) * [Feat] add safety to collections * minor fix	2023-08-11 11:19:26 +08:00
Zaida Zhou	f4c70ba6c3	[Feature] Support filtering specified levels message (#187 ) * Support filtering message * minor fix	2023-08-11 10:46:46 +08:00
Zaida Zhou	f256abffd3	[Enhancement] Skip invalid keys to avoid requesting API (#184 ) * Skip invalid keys to avoid requesting API * get expected key * print warning info	2023-08-10 18:41:43 +08:00
Ma Zerun	59bf56349c	[Feature] Support CUDA_VISIBLE_DEVICES and multiple tasks on one GPU (#148 ) * [Feature] Support CUDA_VISIBLE_DEVICES and multiple tasks on one GPU * Fix UT * Update according to comments	2023-08-10 16:53:03 +08:00
Tong Gao	312095de9d	[Fix] meta template & unit tests (#170 )	2023-08-10 16:49:13 +08:00
liushz	ed248af136	[Fix] Fix some sc errors (#177 ) * Update sc * Update sc doc * Apply suggestions from code review Co-authored-by: Hubert <42952108+yingfhu@users.noreply.github.com> --------- Co-authored-by: liuhongwei <liuhongwei@pjlab.org.cn> Co-authored-by: Hubert <42952108+yingfhu@users.noreply.github.com>	2023-08-10 16:40:32 +08:00
Tong Gao	2931f3dcb8	[Enhancement] Add humaneval postprocessor for GPT models & eval config for GPT4, enhance the original humaneval postprocessor (#129 ) * [Enhancement] Enhance humaneval postprocessor * add human-eval testcase * update * update --------- Co-authored-by: Leymore <zfz-960727@163.com>	2023-08-10 16:31:12 +08:00
Songyang Zhang	3f36db3b06	[Feature] Support turbomind (#166 ) * support turbomind * update doc * Update docs/en/advanced_guides/evaluation_turbomind.md Co-authored-by: Tong Gao <gaotongxiao@gmail.com> * Update docs/zh_cn/advanced_guides/evaluation_turbomind.md Co-authored-by: Tong Gao <gaotongxiao@gmail.com> * Update docs/zh_cn/advanced_guides/evaluation_turbomind.md Co-authored-by: Tong Gao <gaotongxiao@gmail.com> * Update docs/en/advanced_guides/evaluation_turbomind.md Co-authored-by: Tong Gao <gaotongxiao@gmail.com> * update --------- Co-authored-by: Tong Gao <gaotongxiao@gmail.com>	2023-08-10 16:25:11 +08:00
Leymore	e7fc54baf1	[Feature] Add Xiezhi SQuAD2.0 ANLI (#101 ) * add Xiezhi SQuAD2.0 ANLI; update WSC * update * update * update doc string	2023-08-10 14:04:18 +08:00
Yuan Liu	a205629ff3	[Feature]: Refactor input and output (#176 ) * [Feature]: Refactor input and output * [Feature]: Update tasks	2023-08-10 14:01:28 +08:00
Leymore	876ade71a5	[Fix] Fix AGIEval multiple choice (#137 ) * update agieval data * rename variables	2023-08-10 11:38:24 +08:00
Tong Gao	e6194df29e	[Fix] Use a copy of the config object in Task (#174 )	2023-08-09 15:24:49 +08:00
Haodong Duan	d5d4f47371	[API] Refine OpenAI (#175 )	2023-08-09 12:38:57 +08:00
Zaida Zhou	af436f5951	[Feature] Calculate max_out_len without hard code for OpenAI model (#158 ) * calulate max_out_len without hard code * set default value * update configs * Update configs/eval_gpt3.5.py Co-authored-by: Tong Gao <gaotongxiao@gmail.com> --------- Co-authored-by: Tong Gao <gaotongxiao@gmail.com>	2023-08-08 15:16:56 +08:00
Yuan Liu	2f1949e7a1	[Feature]: Add mm suport for local (#169 )	2023-08-08 14:21:58 +08:00
Tong Gao	bbdedc6c95	[Enhancement] Optimize OpenAI models (#128 ) * [Feature] Enhance OpenAI API, add example config for GPT evaluation	2023-08-03 14:55:16 +08:00
Haodong Duan	d17a5b94fa	[Refine] Refine PR #122 (#123 ) * update * update	2023-08-03 14:54:38 +08:00
Yuan Liu	191a3f6f9d	[Feature]: Use multimodal (#73 ) * [Feature]: Add minigpt-4 * [Feature]: Add mm local runner * [Feature]: Add instructblip * [Feature]: Delete redundant file * [Feature]: Delete redundant file * [Feature]: Add README to InstructBLIP * [Feature]: Update MiniGPT-4 * [Fix]: Fix lint * [Feature]add omnibenchmark readme (#49) * add omnibenchmark readme * fix * Update OmniMMBench.md * Update OmniMMBench.md * Update OmniMMBench.md * [Fix]: Refine name (#54) * [Feature]: Unify out and err * [Fix]: Fix lint * [Feature]: Rename to mmbench and change weight path * [Feature]: Delete Omni in instructblip * [Feature]: Check the avaliablity of lavis * [Fix]: Fix lint * [Feature]: Refactor MM * [Refactor]: Refactor path * [Feature]: Delete redundant files * [Refactor]: Delete redundant files --------- Co-authored-by: Wangbo Zhao(黑色枷锁) <56866854+wangbo-zhao@users.noreply.github.com>	2023-08-03 11:07:50 +08:00
Tong Gao	8b163bd8e9	[Feature] Several enhancements (#142 )	2023-08-01 18:19:49 +08:00
Tong Gao	c00179d46b	[Feature] Evaluating acc based on minimum edit distance, update SIQA (#130 ) * [Feature] Support evaluating acc based on minimum edit distance, update SIQA * update	2023-08-01 14:24:27 +08:00
Leymore	d862f570aa	[Feature] Add SC (#126 ) * add self-consistency * add CoT method Self-Consistency * fix typo error and update openicl_eval * add tydiQA-GoldP task * fix sc * rename gsm8k_sc * fix sc * add self-consistency doc * refine sc --------- Authored-by: liushz <qq1791167085@163.com>	2023-07-28 17:29:37 +08:00
Haodong Duan	538b439302	[Fix] Fix seed in HFEvaluator (#122 )	2023-07-28 11:29:01 +08:00
Haodong Duan	46c9645753	[Feature] Allow explicitly setting the temperature for API model (#121 ) * allow explicitly setting the temperature * update	2023-07-28 11:28:15 +08:00
gowithme	57fcfc975a	[Feature] Support intern lanuage model (#51 ) * support internLM * support internLM * simplify intern model files * update storage_manager * support internLM * Modify the file organization structure * support internLM * support internLM * support internLM * support internLM * change some details	2023-07-27 18:49:36 +08:00
Hubert	b7184e9db5	[Refactor] Update crows-pairs evaluation (#98 ) * [Refactor] Update crows-pairs evaluation * [Refactor] Update crows-pairs evaluation * minor	2023-07-26 11:21:32 +08:00
Haonan Li	e9cdb24ddd	[Feature] Add CMMLU dataset (#91 ) * add CMMLU * debug cmmlu * add slurm args `qos` * fix format: space before comment * remove unused variable * change the location of `answer is` --------- Co-authored-by: 李浩楠 <lihaonan@lihaonandeMacBook-Air.local> Co-authored-by: 李浩楠 <haonan.li> Co-authored-by: Leymore <zfz-960727@163.com>	2023-07-25 10:14:27 +08:00
Haodong Duan	6e885d668b	force utf-8 encoding for all non-dataset fileios (#97 )	2023-07-25 10:06:01 +08:00
Leymore	3fe5ee096c	[Feature] Add heuristic size partitioner (#63 ) * [Feature] Add heuristic size partitioner * update	2023-07-20 11:53:24 +08:00
Leymore	eea8b04417	[Feature] Add llama-2 models (#81 ) * add llama-2 models * update docs --------- Co-authored-by: gaotongxiao <gaotongxiao@gmail.com>	2023-07-19 19:51:29 +08:00
Hubert	f83e125e5a	[Feat] Support CValues Responsibility dataset (#78 ) * [Feat] support CValues * minor fix	2023-07-18 18:45:15 +08:00
LZH	26e2f171f4	[Feature] Support load PEFT adapter for HuggingFace model (#74 ) * support peft for HuggingFace model * add docstring	2023-07-18 16:21:43 +08:00
liushz	f36c0496f3	[Feature] Add tydiqa-goldp (#75 ) Co-authored-by: liuhongwei <liuhongwei@pjlab.org.cn>	2023-07-18 14:54:35 +08:00
Tong Gao	311bf0daa7	[Fix] Fix CI (#70 ) * [Fix] Fix CI * [Fix] Fix CI * [Fix] Fix CI * update	2023-07-17 19:10:59 +08:00
Tong Gao	29006e39c0	[Fix] Fix circular import of PromptTemplate (#71 )	2023-07-17 19:09:38 +08:00
Tong Gao	1e44541730	[Enhancement] Test linting in CI and fix existing linting errors (#69 ) * [Enhancement] Test linting in CI * fix linting	2023-07-17 15:59:10 +08:00
Leymore	1326aff77e	[Feature] Add logger info and remove dataset bugs (#61 ) * Add logger info and remove dataset bugs * fix typo	2023-07-17 14:26:30 +08:00
Tong Gao	7ee5a86fee	[Feature] Enhance OpenAI API, add example config for GPT evaluation (#53 ) * [Feature] Enhance OpenAI API, add example config for GPT evaluation * fix	2023-07-12 16:43:46 +08:00
Hubert	f5103f93dd	[Feat] add bs for perspective api eval (#50 ) * [Feat] add bs for perspective api eval * fix according to comments * fix according to comments	2023-07-12 16:26:01 +08:00
Hubert	c8f1d513b2	[Fix] fix clp inferencer (#44 )	2023-07-11 14:54:39 +08:00
Tong Gao	0625294e5f	[Fix] Fix OpenICLInferTask (#41 )	2023-07-10 16:12:01 +08:00
Ma Zerun	805293a9f2	Auto re-generate port number during retry (#24 ) * Auto re-generate port number during retry * Fix slurm command	2023-07-07 17:25:56 +08:00
Hubert	7f8eee4725	[Docs] add en docs (#15 ) * add en docs * update --------- Co-authored-by: gaotongxiao <gaotongxiao@gmail.com>	2023-07-06 12:58:44 +08:00
Leymore	86d5ec3d0f	Update configs (#9 ) * Update implements * Update	2023-07-06 12:27:41 +08:00
Hubert	5c19c8c5fc	[Docs] add issue and pr template (#12 ) * [Feat] add issue and pr template * minor add utils * minor fix	2023-07-06 11:55:01 +08:00
Tong Gao	986a44cedd	Tranlate lark messages (#5 )	2023-07-05 18:40:05 +08:00
Tong Gao	719ba34d1b	[Enhancement] Update prompt hash computation (#2 )	2023-07-05 18:29:07 +08:00
Ma Zerun	5840c7655c	Update start guide (#4 )	2023-07-05 18:26:26 +08:00
yuzhaohui	dcf11cf8fd	New logo and update setup.py	2023-07-05 06:54:06 +00:00
mzr1996	04dd01a235	Update configs and code	2023-07-05 11:45:08 +08:00
Leymore	c94cc94348	Add release contribution	2023-07-05 03:15:31 +00:00
tonysy	e6b5bdcb87	OpenCompass Public MR	2023-07-05 03:15:21 +00:00
Ezra-Yu	cbe9fe2cdb	Add Release Contraibution	2023-07-05 02:22:40 +00:00
cky	36f111100f	update datasets	2023-07-05 01:45:26 +00:00
mzr1996	3cfe73de3f	Support a batch of datasets.	2023-07-05 01:30:27 +00:00
kennymckormick	78478e961e	[Code] Update opencompass/datasets/agieval/__init__.py	2023-07-05 00:28:07 +00:00
yingfhu	fb11108723	[Feat] support opencompass	2023-07-04 22:11:33 +08:00
gaotongxiao	7d346000bb	initial commit	2023-07-04 21:34:55 +08:00

... 5 6 7 8 9 ...

609 Commits