0% found this document useful (0 votes)
23 views8 pages

Numeric Sled Project Task Guide

Uploaded by

habibullasardar
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
23 views8 pages

Numeric Sled Project Task Guide

Uploaded by

habibullasardar
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Numeric Sled Project Instructions

Introduction

The tasks you will see are composed of three components: prompt writing, rating two model
responses, preference ranking, and justifications. Each component should be thought of as a
unique skill that you should develop. However, they’re all working at the same outcome, which is
producing and evaluating model responses.

This document will walk you over everything you need to keep in mind while tasking to ensure
you produce high quality tasks.

🖊️ Part 1: Writing High Quality Prompts


Recall that prompts are composed of three things:
1. Context
2. Instruction (The Main Request)
3. Constraints

The main request dictates what the model should discuss. The constraints dictate the degree of
freedom the model has when writing its response.

A good prompt is:


1. Specific
2. Detailed
3. Practical
4. Has at least 1 constraint
A strong constraint explicitly and clearly defines something non-trivial that the AI must or must
not do during the conversation, and makes the AI work hard to comply. See types of constraints
you can use in a prompt here. Every prompt is required to have constraints. These are not the
same thing as just giving the AI an instruction.
An instruction tells the AI what to do.
A constraint limits the AI’s degrees of freedom in some way.

Examples:

● “Tell me about World War I.”


○ No constraint.
● “Tell me about WW1. Do not mention France.”
○ One constraint.
● “Tell me about WW1. Do not mention France or Germany.”
○ One constraint. The constraint has multiple elements but is still one constraint.
● “Tell me about WW1. Do not mention France. Respond as a professor of military history
who loves to talk about artillery.”
○ Two constraints.
● “Write me a poem about WW1 in Shakespearean sonnet form. I may then ask you a
follow-up question about France, which you should answer.”
○ One constraint. Sonnet form is the constraint. Notifying the AI that you will ask
more questions does not constrain the AI’s responses.

Note: The prompts need to have a minimum of 15 words and a maximum of 200 words. Avoid
using very long text spans, including Reference text.

🪣 Part 2: Performing Ratings


The next step is to rate the two model responses you will see based on three criteria: Accuracy,
Instruction Following, and Writing Quality.
🎯Accuracy

​ Rate the response a 1: The response contains false information or unverifiable


claims.
​ Rate the response a 2: The response contains minor inaccuracies that do not
majorly affect the main request.
​ Rate the response a 3: The response is mostly accurate but may include
hard-to-verify information.
​ Rate the response a 4: The response is accurate with only minor hard-to-verify
statements that do not impact the overall correctness.
​ Rate the response a 5: The response is wholly accurate and verifiable through
trusted
​ sources.
🔬Instruction Following

​ Rate the response a 1: The response fails to address the main request of the
prompt and all of its constraints.
​ Rate the response a 2: The response addresses the main request but misses
important details or constraints.
​ Rate the response a 3: The response addresses the main request and all
constraints, though it may include some unnecessary information.
​ Rate the response a 4: The response fully addresses the main request and all
constraints and includes minimal unnecessary information.
​ Rate the response a 5: The response fully addresses the main request and all
constraints, and every part of the response is relevant to the prompt's requests.

✍️Writing Quality
​ Rate the response a 1: The response has many spelling or grammatical errors,
is disorganized, and repeats ideas without intention. The tone is inappropriate.
​ Rate the response a 2: The response has several errors impacting readability, is
somewhat disorganized, or explains concepts inefficiently. The tone is
acceptable.
​ Rate the response a 3: The response is clunky or suboptimally worded but is
understandable, organized, and explains concepts most efficiently. The tone is
appropriate.
​ Rate the response a 4: The response is well-organized with smooth transitions,
communicates concepts efficiently, and the tone is contextually appropriate and
professional.
​ Rate the response a 5: The response is exceptionally well-organized with
seamless transitions, communicates concepts with precision and clarity, and the
tone perfectly suits the context, enhancing the overall quality of the response.

Response Rating Rubric - Combined


Criteria 1 [Terrible] 2 [Poor] 3 [Okay] 4 [Good] 5 [Amazing]

Instruction The response The response The response The response fully The response
Following fails to address addresses the addresses the addresses the main fully addresses
the main request main request but main request and request and all the main request
of the prompt misses important all constraints, constraints and and all
details or though it may includes minimal constraints, and
constraints. include some unnecessary every part of the
unnecessary information. response is
information. relevant to the
prompt's
requests.

Accuracy The response The response The response is The response is The response is
contains false contains minor mostly accurate accurate with only wholly accurate
information or inaccuracies that but may include minor hard-to-verify and verifiable
unverifiable do not majorly hard-to-verify statements that do through trusted
claims. affect the main information. not impact the sources.
request. overall correctness.

Writing Quality The response The response The response is The response is The response is
has many has several clunky or well-organized with exceptionally
spelling or errors impacting suboptimally smooth transitions, well-organized
grammatical readability, is worded but is communicates with seamless
errors, is somewhat understandable, concepts efficiently, transitions,
disorganized, disorganized, or organized, and and the tone is communicates
and repeats explains explains concepts contextually concepts with
ideas without concepts most efficiently. appropriate and precision and
intention. The inefficiently. The The tone is professional. clarity, and the
tone is tone is appropriate. tone perfectly
inappropriate. acceptable. suits the context,
enhancing the
overall quality of
the response.

🔋Overall Score
The overall score should be determined by taking into account these response characteristics:
accuracy, instructions following, safety, and formatting/writing quality. Think about the overall
usefulness of the response before choosing a score. Do not only consider accuracy and
instruction following.
⚖️ Part 3: Comparing the two responses
🏅Preference Ranking
After rating each response on the 4 dimensions, rank your preference of the responses on a 1-5
Likert scale.

1 2 3 4 5

Response A is Response A is Response B is Response B is


Neutral
much better slightly better slightly better much better

Select 1 or 5 if:
● One response is (almost) completely flawless, whereas the other response has
major issues.
● Both responses have issues, but one response is substantially better.
● One response is completely flawless, whereas the other response has minor
issues.
Select 2 or 4 if:
● There are small differences in writing quality between the responses
Select 3 if:
● Both responses have major issues and are, therefore, bad (Note: this still applies
if Response A is severely worse than Response B but also has major issues).
● Both responses are identical or near identical in content.

📝Writing Justifications
The final step of the task is to write a justification that explains the reasoning behind your ratings
and preference ranking. While writing a justification, keep the following in mind:

● Start with the conclusion: ex. Response A is better than Response B.


● Add supporting claims: every claim that a justification makes should be grounded in an
rating dimension and supported by an excerpt from the response or a direct reference to
it. Examples:
○ Response A is more accurate than Response B since Response B claims that
[XYZ] which cannot be verified
○ Response A follows instructions better than Response B since it adheres to the
word limit while Response B does not.
○ Notice that these examples are grounded in one of the 4 rating dimensions and
justified by an excerpt from the text or a direct reference to it.
● Be exhaustive: to be sure your justification is exhaustive, call out every dimension where
the 2 responses are scored differently and add excerpts from the responses.
○ Your justification should directly call out any dimension where there is a
difference in scoring across dimensions
■ Think about it: if you rated one response as a 3 on accuracy and the other
as a 4 on accuracy, you must’ve had a reason right? Let us know in the
justification!
○ Your justification should also call out any dimension where both responses are
rated as a fail (ie 1 or 2)
■ This is helpful context for the model developers. If both responses have
bad writing quality - let them know in the justification.

Example Justification: Response A is much better than Response B. Response B is better at


instruction following than Response A since it provides three directors as requested while
Response A provides 5. Despite this, Response B has a major accuracy error. Response B
states that Quentin Tarantino’s Pulp Fiction (1994) grossed over $1B, however, Pulp Fiction has
only grossed $212,891,760 worldwide. This critical factuality error makes Response A the better
of the two responses.

Structure your justification as follows:



1. Start with your conclusion (e.g. “Response A is better than Response B”)
2. Add supporting claims: every claim that a justification makes should be grounded in a
rating dimension and supported by an excerpt from the response or a direct reference to
it
3. Be exhaustive: be sure your justification is exhaustive, call out every dimension where
the 2 responses are scored differently, and add excerpts from the responses

Appendix

Types of Constraints:

Different types of constraints you can use in a prompt:

Content ● Request that the model include or exclude


specific information.

Language Style ● Specify any linguistic guidelines, such as


avoiding jargon, using plain language, or
adapting the response for a specific audience.
● Requesting the use of specific literary devices
or literary styles.
Guidelines for ● Teach the model how to handle situations
Dealing with where it doesn't have a definitive answer.
Uncertainty Instruct it to admit uncertainty or offer
educated guesses rather than providing
inaccurate information

Bias and ● Instruct the model to correct any biased


Sensitivity language or assumptions in the user's input.

Engagement and ● Instruct the model to ask clarifying questions if


Interaction the user's input is ambiguous or unclear.
● Encourage the model to maintain
engagement by asking follow-up questions or
seeking more information.

Length ● Limit the length of responses or prescribe a


minimum length of response.

Geographic and ● Instruct the model to consider geographic or


Location-Based location-based context when generating
responses. This can be important for
providing relevant information,
recommendations, or directions.

Expertise Level ● Specify the level of expertise the model


should assume when responding. e.g. it can
act as a beginner, intermediate, or expert in a
particular field

Time Sensitivity ● Instruct the model to provide responses that


are relevant to a particular time frame or
historical context

Humor and Wit ● Guide the model on when and how to use
humor, wit, or clever wordplay in responses

Formatting and ​ Instruct the model on how to structure its


Structure responses, such as using bullet points for
lists, headers for sections, tables, LaTex, etc.

You might also like