Prometheus-7B-v2.0 Open-Source Language Model - Alternative to GPT-4 for Fine-Grained Evaluation Selection

Prometheus 7b V2.0

Developed by prometheus-eval

Prometheus 2 is a language model based on Mistral-Instruct, specifically designed for fine-grained evaluation and reinforcement learning from human feedback, serving as an alternative to GPT-4 evaluation.

Large Language Model

Transformers

EnglishOpen Source License:Apache-2.0 #Language Model Evaluation #Feedback Generation #RLHF Reward Model

Downloads 13.07k

Release Time : 2/13/2024

Model Overview

This model supports both absolute and relative rating evaluation methods, enhancing performance through weight merging techniques, suitable for evaluating content generated by language models.

Model Features

Dual-Mode Evaluation

Supports both absolute rating (direct evaluation) and relative rating (pairwise ranking) evaluation modes

Weight Merging Technique

Enhances performance in each rating format through innovative weight merging methods

Fine-Grained Feedback

Capable of generating detailed quality feedback and comparative analysis, not just simple ratings

Model Capabilities

Text Generation

Quality Evaluation

Feedback Generation

Pairwise Comparison

Use Cases

Language Model Evaluation

Generated Content Quality Evaluation

Evaluates the quality of content generated by language models and provides detailed feedback

Can serve as an alternative to GPT-4 for automated evaluation

Model Comparative Evaluation

Compares the relative quality of outputs from two different models

Provides objective comparative analysis

Reinforcement Learning

RLHF Reward Model

Acts as a reward model in reinforcement learning from human feedback

Delivers fine-grained feedback signals

🚀 Prometheus 2

Prometheus 2 serves as an alternative to GPT - 4 for fine - grained evaluation of underlying LLMs and as a reward model for RLHF.

🚀 Quick Start

Prometheus 2 can be used for fine - grained evaluation of an underlying LLM and as a Reward model for Reinforcement Learning from Human Feedback (RLHF). It is a language model based on Mistral - Instruct, fine - tuned on 100K feedback from the Feedback Collection and 200K feedback from the Preference Collection. It also supports both absolute grading and relative grading through weight merging.

plot

✨ Features

Alternative to GPT - 4: Prometheus 2 can be used as an alternative for fine - grained evaluation when GPT - 4 is not applicable.
Dual - grading Support: It supports both absolute grading (direct assessment) and relative grading (pairwise ranking) through weight merging, and the performance on each format is also improved.
Fine - tuned on Large - scale Data: It is fine - tuned on large - scale feedback data from the Feedback Collection and Preference Collection.

📚 Documentation

Links for Reference

Homepage: In Progress
Repository: https://github.com/prometheus-eval/prometheus-eval
Paper: https://arxiv.org/abs/2405.01535
Point of Contact: seungone@cmu.edu

Model Details

Model Description

Property	Details
Model Type	Language model
Language(s) (NLP)	English
License	Apache 2.0
Related Models	All Prometheus Checkpoints
Resources for more information	Research paper, GitHub Repo

Prometheus is trained in two different sizes (7B and 8x7B). You can check the 8x7B sized LM on this page. Also, check out our dataset on this page and this page.

Prompt Format

Basic Usage

Prometheus 2 has different prompt formats for absolute grading and relative grading.

Absolute Grading (Direct Assessment) Prometheus requires 4 components in the input: An instruction, a response to evaluate, a score rubric, and a reference answer.

###Task Description:
An instruction (might include an Input inside it), a response to evaluate, a reference answer that gets a score of 5, and a score rubric representing a evaluation criteria are given.
1. Write a detailed feedback that assess the quality of the response strictly based on the given score rubric, not evaluating in general.
2. After writing a feedback, write a score that is an integer between 1 and 5. You should refer to the score rubric.
3. The output format should look as follows: \"Feedback: (write a feedback for criteria) [RESULT] (an integer number between 1 and 5)\"
4. Please do not generate any other opening, closing, and explanations.

###The instruction to evaluate:
{orig_instruction}

###Response to evaluate:
{orig_response}

###Reference Answer (Score 5):
{orig_reference_answer}

###Score Rubrics:
[{orig_criteria}]
Score 1: {orig_score1_description}
Score 2: {orig_score2_description}
Score 3: {orig_score3_description}
Score 4: {orig_score4_description}
Score 5: {orig_score5_description}

###Feedback:

After this, you should apply the conversation template of Mistral. You can find the conversation class at this link.

conv = get_conv_template("mistral")
conv.set_system_message("You are a fair judge assistant tasked with providing clear, objective feedback based on specific criteria, ensuring each assessment reflects the absolute standards set for performance.")
conv.append_message(conv.roles[0], dialogs['instruction'])
conv.append_message(conv.roles[1], None)
prompt = conv.get_prompt()

x = tokenizer(prompt,truncation=False)

As a result, a feedback and score decision will be generated, divided by a separating phrase [RESULT]

Relative Grading (Pairwise Ranking) Prometheus requires 4 components in the input: An instruction, 2 responses to evaluate, a score rubric, and a reference answer.

###Task Description:
An instruction (might include an Input inside it), two responses to evaluate (denoted as Response A and Response B), a reference answer, and an evaluation criteria are given.
1. Write a detailed feedback that assess the quality of the two responses strictly based on the given evaluation criteria, not evaluating in general.
2. Make comparisons between Response A, Response B, and the Reference Answer. Instead of examining Response A and Response B separately, go straight to the point and mention about the commonalities and differences between them.
3. After writing the feedback, indicate the better response, either "A" or "B".
4. The output format should look as follows: "Feedback: (write a feedback for criteria) [RESULT] (Either "A" or "B")"
5. Please do not generate any other opening, closing, and explanations.

###Instruction:
{orig_instruction}

###Response A:
{orig_response_A}

###Response B:
{orig_response_B}

###Reference Answer:
{orig_reference_answer}

###Score Rubric:
{orig_criteria}

###Feedback:

After this, you should apply the conversation template of Mistral. You can find the conversation class at this link.

conv = get_conv_template("mistral")
conv.set_system_message("You are a fair judge assistant assigned to deliver insightful feedback that compares individual performances, highlighting how each stands relative to others within the same cohort.")
conv.append_message(conv.roles[0], dialogs['instruction'])
conv.append_message(conv.roles[1], None)
prompt = conv.get_prompt()

x = tokenizer(prompt,truncation=False)

As a result, a feedback and score decision will be generated, divided by a separating phrase [RESULT]

License

Feedback Collection, Preference Collection, and Prometheus 2 are subject to OpenAI's Terms of Use for the generated data. If you suspect any violations, please reach out to us.

Citation

If you find the following model helpful, please consider citing our paper!

BibTeX:

@misc{kim2023prometheus,
    title={Prometheus: Inducing Fine-grained Evaluation Capability in Language Models},
    author={Seungone Kim and Jamin Shin and Yejin Cho and Joel Jang and Shayne Longpre and Hwaran Lee and Sangdoo Yun and Seongjin Shin and Sungdong Kim and James Thorne and Minjoon Seo},
    year={2023},
    eprint={2310.08491},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}

@misc{kim2024prometheus,
    title={Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models},
    author={Seungone Kim and Juyoung Suk and Shayne Longpre and Bill Yuchen Lin and Jamin Shin and Sean Welleck and Graham Neubig and Moontae Lee and Kyungjae Lee and Minjoon Seo},
    year={2024},
    eprint={2405.01535},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}

Featured Recommended AI Models

Empowering the Future, Your AI Solution Knowledge Base

English 简体中文繁體中文にほんご