cai/benchmarks/seceval/eval/README.md

3.1 KiB

Evaluating on SecEval Dataset

Install Dependencies

pip install -r requirements.txt

Download Dataset

You can download the json file of the dataset by running.

wget https://huggingface.co/datasets/XuanwuAI/SecEval/blob/main/questions.json

Evaluate on a Model

Usage

usage: eval.py [-h] [-o OUTPUT_DIR] -d DATASET_FILE [-c] [-b BATCH_SIZE] -B {remote_hf,azure,textgen,local_hf} -m MODELS [MODELS ...]

SecEval Evaluation CLI

options:
  -h, --help            show this help message and exit
  -o OUTPUT_DIR, --output_dir OUTPUT_DIR
                        Specify the output directory.
  -d DATASET_FILE, --dataset_file DATASET_FILE
                        Specify the dataset file to evaluate on.
  -c, --chat            Evaluate on chat model.
  -b BATCH_SIZE, --batch_size BATCH_SIZE
                        Specify the batch size.
  -B {remote_hf,azure,textgen,local_hf}, --backend {remote_hf,azure,textgen,local_hf}
                        Specify the llm type. remote_hf: remote huggingface model backed, azure: azure openai model, textgen: textgen backend, local_hf: local huggingface model backed
  -m MODELS [MODELS ...], --models MODELS [MODELS ...]
                        Specify the models.

Select a Backend

There are four backends supported for evaluating on SecEval dataset. Some backends require additional environment variables to be set.

Local Huggingface Model Backend

Local Huggingface Model Backend is used for evaluating on models that are stored locally. You can download the models from Huggingface and use this backend to evalute the model.

export HF_LOCAL_MODEL_DIR=<path to local huggingface model>

Remote Huggingface Model Backend

Remote Huggingface Model Backend is used for evaluating on models that are stored on Huggingface. You can use this backend to evalute the model.

Azure OpenAI Model Backend

Azure OpenAI Model Backend is used for evaluating on models that are deployed on Azure OpenAI.

export OPENAI_API_KEY=<your openai api key>
export OPENAI_API_ENDPOINT=<your openai api endpoint>

text-generation-webui Backend

text-generation-webui is a service used for deploying and testing models. We have supported evaluating on models deployed on text-generation-webui.


export TEXTGEN_MODEL_URL=<your textgenerationwebui api url>

Run Evaluation

If you want to evaluate on a base model, you can run the following command:

python eval.py -d <path to dataset file> -B <backend> -m <model name> -o <output directory>

If you want to evaluate on a chat model such as gpt-3.5-turbo, you should add the -c flag:

python eval.py -d <path to dataset file> -B <backend> -m <model name> -o <output directory> -c

Inspection the Results

The results will be saved in the output directory. The result is a JSONObject contains following field

{
    "score_fraction": {
        "topic_name": "topic_score_fraction",
    },
    "score_float": {
        "topic_name": "topic_score_float",
    },
    "detail": ["detail1", "detail2",] 
}