0% found this document useful (0 votes)
62 views6 pages

Task Segment Labeling Guidelines

Uploaded by

apjame
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as ODT, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
62 views6 pages

Task Segment Labeling Guidelines

Uploaded by

apjame
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as ODT, PDF, TXT or read online on Scribd

Introduction & Expectations

Welcome to Task Segment Labeling

Exit

Goal

The goal of task description text labels is to provide a high level description of tasks that
is factually correct to what is occurring visually. Labels are machine annotated and human
sample reviewed.

Expectations

Minimum Pace

1x realtime — review 1 hour of footage per 1 hour of work

Videos play at 4x speed by default to help you work efficiently

Accuracy Target

97% correct labels and edits

Work will be audited and reviewed for correct review or edit of labels

Performance Review: A performance review and evaluation will occur if work falls below or
exceeds expectations.

When to Reject Annotations

Annotations should be FAILED if any of these apply

Exit

Important: If you see any of the following issues, the annotation must be rejected.

Hallucinations

Labels are annotating hallucinations — the label did not happen at all. The action or object
described in the label is completely fabricated.

Incorrect Timestamps

The timestamps are incorrect with >2 seconds of error. The timing is significantly off from when
the action actually occurs.

Overlapping Annotations
Annotations are overlapping one another — timeframes bleed into the other. Each segment
should have distinct, non-overlapping time boundaries.

When to Reject Annotations

Annotations should be FAILED if any of these apply

Exit

Important: If you see any of the following issues, the annotation must be rejected.

Hallucinations

Labels are annotating hallucinations — the label did not happen at all. The action or object
described in the label is completely fabricated.

Incorrect Timestamps

The timestamps are incorrect with >2 seconds of error. The timing is significantly off from when
the action actually occurs.

Overlapping Annotations

Annotations are overlapping one another — timeframes bleed into the other. Each segment
should have distinct, non-overlapping time boundaries.

When to Validate Annotations

Annotations should be PASSED if these criteria are met

Exit

Good annotations meet all of the following criteria.

Factual Description

The task description is factual — it accurately describes what is happening in the video.

Appropriately Broad

The task description is factual and broad enough to encompass the purpose of all tasks in the
segment.

Minor Timestamp Tolerance

Minor timestamp differences are acceptable — <2 seconds of error

General Labeling Rules

Standards and requirements for all labels


Exit

Objective: Create a 100% factual timeline of actions for head-mounted camera footage.

Labeling Standards

Action Style

Labels should describe the actions in the video.

Inactivity

Label any period without specific task activity as "no action"

Factuality

Strictly describe visible task goals only. Do not make assumptions, guesses, or infer intent.

Object Details

Omit descriptive attributes such as brands, colors, or models if they are uncertain.

Example:black iPhone→phone

Format Preference

Use simple, descriptive task labels:

Using scissors to strip insulation from the black cable

pick up cable

Label Rules

Duration

Minimum: 8 seconds
Maximum: 40 seconds per label

Coverage

100% of the video must be annotated with no temporal gaps.

Output Format

Line format:

mm:ss - mm:ss, label_text

Examples: Hallucinations & Going Coarse


Learn when to reject and how to fix incorrect labels

Exit

Hallucinating Actions and Objects

Problem: Incorrect object, no flour used. Incorrect action, no pouring.


Choice: Reject

Uncertain Rotation → Go Coarse

When you're unsure about the specific action direction, use a more general (coarse) term.

Episode ID: 6901b06ea95c8e2275bf8ffb

Problem: User is not removing the screws. They seem to be installing them.

Solution: Delete "removing" and reword to "rotating" or "fixing" which is more coarse and
factual.

Old Label:

Removing screws from the top edge of the phone chassis

New Label:

Rotating screws from the top edge of the phone chassis

Episode ID: 6923fffee3d0b310ce4e3354

Problem: User is removing the screws, not tightening.

Solution: Reword "reassembling" to "removing" and "tightening" to "removing".

Old Label:

Reassembling the socket components and tightening the main screw

New Label:

Removing the socket components and removing the main screw

Examples: Word Flexibility & Granular Details

Understanding when labels are acceptable and when details matter

Exit

Word Choices Are Flexible

Different word choices can be equally correct as long as they accurately convey the action.
Episode ID: 68c6742b70b739c3b9948b58

placing the bouquet of flowers diagonally onto the center of the white wrapping paper

Explanation: With flowers laid out front, the recorder moves the flowers from the left side to
the middle of the wrapping paper. Here "placing" is enough to convey that they are "moving" or
"picking" and "placing" the flowers onto the wrapping paper.

Episode ID: 68c4c83f4488c73eb899ef5e

adjusting_pink_string

Explanation: In the few second label, the recorder picks up the spool of string and is adjusting
the string to prepare to pull and cut it. The label "adjust_pink_string" is still correct.

Incorrect Granular Details

If the annotation describes something specific about an object (brand, color, quantity) AND it's
incorrect — reject the annotation.

"Pick up two spoons" but there was only 1 spoon → Reject and edit "two" to "one"

"Place the nike shirt down" but not a Nike shirt → Reject and delete "nike"

"Roll the second egg roll" and there's a pile of egg rolls on table → Validate (sequential counting
is okay)

Examples: Half True Labels & Counting

Final examples and confirmation

Exit

Half True Labels and Multiple Actions

Labels that mention two or more actions — if one of the actions is untrue and not actually
done — should be rejected.

Episode ID: 68c3eaddfa7cedbd1644ac88

Old Label:

pick_up_weight_and_frame

New Label:

Picking up weight
Explanation: The label mentions picking up both weight and frame, however, only the weight
was interacted with in the video. This is half false so reject and edit.

Episode ID: 69063f08b36f746e6646ed1a

Placing a label in a new plastic carton and starting to fill it with eggs

Explanation: The label mentions placing a label which is correct. But it is incorrect that they fill
it with eggs. Reject this label.

Sequential Counting or Context

In coarse labeling, we only describe the action being done. We don't keep track of which exact
object is which.

"Potting the second seedling into a black polybag" but not necessarily the second seedling
→ Validate

Reference: Good Video Example

Episode ID: 690824af4207384788b95bd4

Common questions

Powered by AI

The significance of having a minimum (8 seconds) and maximum (40 seconds) label duration is centered on capturing the key actions within a consistent timeframe, which aids in preserving the context and flow of actions without losing essential details . This defined duration framework affects video annotation accuracy by preventing overlapped or excessively lengthy segments that could introduce errors, ensuring that annotations reflect a concise and factual execution of actions within the video.

Labeling tasks must ensure 100% of the video is annotated without temporal gaps to provide a complete representation of the visual content . Full coverage is necessary because it ensures that no actions are missed, which is crucial for creating a reliably detailed timeline that reflects every moment in the footage, thereby aiding in the development of comprehensive algorithms.

Annotation output should adhere to a specific format such as mm:ss - mm:ss, label_text to standardize the data entry process, making it easier to parse and analyze by machines . This uniformity plays a critical role in task segment labeling efficiency as it ensures consistency, minimizes human error in data entry, and facilitates smooth integration with automated systems that rely on precise inputs for accurate processing.

Word flexibility should be applied when different word choices equally convey the action being performed without altering the factual accuracy . This flexibility entails using descriptive terms that best fit the action observed, even if they differ slightly from the initial label, as long as they maintain the integrity of the action described. This allows for diverse descriptions and can accommodate a range of synonymous terms that correctly express the visible actions.

Annotations should be PASSED if they meet the criteria of providing a factual description, being appropriately broad to encompass the purpose of all tasks in the segment, and having minor timestamp differences of less than 2 seconds . These criteria ensure that annotations accurately reflect the visual content and context without unnecessary specificity, which is crucial for maintaining a high accuracy target of 97% and ensuring reliable machine learning training data.

It is important to avoid hallucinations, which are labels that describe actions or objects not actually present, because these can introduce inaccuracies that compromise the data's utility in machine learning models . If hallucinations occur, annotations must be rejected because they falsely represent the visual data, leading to erroneous algorithmic learning and potentially ineffective outcomes in model performance.

Word choice flexibility should be applied by selecting terms that accurately convey the action without deviating from factual representation . This means using synonymous, descriptive words that portray the action observed, such as 'placing' being used to convey movement, ensuring the label accurately reflects the visual action while accommodating variations in language. This flexibility helps maintain the descriptive yet factual nature of annotations.

Understanding when to reject 'half-true' labels, which mention multiple actions where only some occur, prevents inaccuracies from compromising the data set . Rejecting these ensures that annotations are fully aligned with visual truth, preventing misleading data that could affect the performance of algorithms trained on this data. This helps uphold the accuracy standard and target of 97% correctness, crucial for effective and reliable machine intelligence.

Granular details refer to specific information about objects, such as their brand, color, or quantity, in annotations . Incorrect granular details should be handled by rejecting and editing the annotation if the description is wrong, as specific inaccuracies can lead to misleading information that disrupts the factual representation of the scene. This careful editing ensures that only non-error-prone data is used for machine learning purposes.

Adjusting task labels to be factual and coarse rather than inferred is necessary to eliminate assumptions that could inject subjectivity into the data . Maintaining factual and broad labels ensures that only the visible actions are described, which is essential for producing reliable training data for machine learning. Labeling actions coarsely avoids unnecessary specificity that cannot be assured, thus maintaining the integrity of the data.

You might also like