0% found this document useful (0 votes)
5 views6 pages

SideCash Work Platform Reference Guide

This document provides comprehensive guidelines for video captioning tasks on a work platform, detailing the workflow, writing high-level and detailed captions, segmenting videos, and evaluating performance. It emphasizes the importance of accuracy, efficiency, and adherence to specific rules for captioning and segmentation. Annotators are expected to complete tasks within a set timeframe while maintaining a high accuracy score and are warned against cheating.

Uploaded by

tejas nandani
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
5 views6 pages

SideCash Work Platform Reference Guide

This document provides comprehensive guidelines for video captioning tasks on a work platform, detailing the workflow, writing high-level and detailed captions, segmenting videos, and evaluating performance. It emphasizes the importance of accuracy, efficiency, and adherence to specific rules for captioning and segmentation. Annotators are expected to complete tasks within a set timeframe while maintaining a high accuracy score and are warned against cheating.

Uploaded by

tejas nandani
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Work Platform · Video Captioning

Annotation Guidelines
Everything you need to complete video-captioning tasks on work platform.
Designed to be read alongside the tutorial video.

1. What Youʼre Doing and Why


Each task gives you a short video of a real-world task being performed — preparing food, assembling
parts, packing items. Your job is to describe what happens, in two layers:
● High-level caption — one or two plain sentences summarizing the whole video and what task is
being performed.
● Detailed captions — a short caption for each action (segment), describing the specific action
happening in that moment.
Together these teach an AI model both the goal of a task and the step-by-step actions that achieve it.
The model learns from exactly what you write, so precise, literal descriptions matter.

Why you write the high-level caption first


Before captioning individual segments, you write the high-level caption. It feeds our pre-captioning
agent, which makes your detailed captions more accurate and consistent.
That is why the workflow starts by previewing the whole video quickly and writing the high-level
caption before you touch any segments.

The pre-captioning agent drafts — you decide


The agent is a starting point, not the final answer. You segment the video, the agent proposes captions
for each segment, and you check every segment captions against what you actually see—keeping what
is right and fixing what is wrong. Your judgment is the quality bar.

2. The Workflow, Start to Finish


Four stages. Aim to finish a one-minute video in about 5–6 minutes on your first try; you are expected to
get faster as you become familiar.
Stage 1 — Preview and write the high-level caption. Play the video at 8× speed (Shift + ↑) and watch
it through once. While it plays, write the high-level caption.

Stage 2 — Segment the video. Slow the video back to normal speed (Shift + ↓). Create a segment at
every distinct action: pause at the point you want to segment, adjust the scrubber precisely with the
arrow keys, then press C to mark the range. Segments must cover the whole video with no gaps and no
overlaps (Section 4). The tutorial video is the best way to see segmentation done in practice.

Stage 3 — Caption each segment. The pre-captioning agent generates 3 captions for each segment
as you create it. Choose the most correct one and edit as needed. For repetitive segments (e.g.,
scrubbing the floor, a break, then scrubbing again), reuse captions from the Label Bank (Ctrl + B)
instead of retyping (Section 6).
Stage 4 — Review. Play back through the final segments with captions shown and confirm every action
is covered and every caption matches what is on screen. The entire video should be segmented and
captioned.

Time budget (one-minute video)


Stage Target
Preview + high-level caption + first 4–5 ~1 minute
segments
Remaining segments ~1 minute
Caption editing / corrections ~2 minutes
Buffer ~2 minutes
Approx Total Time 5-6 minutes

3. Writing the High-Level Caption


The high-level caption answers one question: what task is happening in this video? Keep it to one or
two sentences of simple, natural English. Describe the overall activity and its goal, not the individual
movements.

Rules
● Use simple English — plain, natural words
● Describe the task — focus on the overall activity and its goal
● Keep it short — one or two sentences
● Do not mention (high-level only) — operator, worker, person, left or right hand, arms, or robotic
arms. Keep the focus on the task, not the body performing it. Avoid heavy technical or machinery
terms unless there is no simpler word.
Hands are handled differently in the two caption types: high-level captions never mention hands; detailed
captions do name the working hand (Section 5). This is intentional, not a contradiction.

Examples
✓ “Coffee is prepared using the coffee machine and poured into a cup.ˮ
✓ “Vegetables are washed and placed into containers.ˮ

✗ “The operator uses the right hand to place the metal object.“
Why: names the operator and a hand; too detailed for a high-level caption.

✗ “The robotic manipulator performs tray alignment.ˮ


Why: too technical; not simple English.

4. Segmenting the Video


Break the video into segments. Each segment is one clear action, or two closely connected actions that
naturally belong together.
Rules
● Cover every action — every meaningful action in the video belongs to a segment
● No gaps — do not leave unannotated time between segments
● No overlaps — two segments must never share the same moment
● How many — a typical one-minute video needs about 10–12 segments, but it depends on the video.
Add as many as it takes to cover every action.
● Combine small connected actions — when two actions naturally go together, cover them in one
segment to avoid over-segmenting

Examples
Timing Verdict
✓ Correct 0.00–3.20 → 3.20–6.10 → Continuous — no gaps or
6.10–9.00 overlaps
✗ Wrong (gap) 0.00–3.20 → 4.00–6.10 Nothing annotated between
3.20 and 4.00
✗ Wrong (overlap) 0.00–3.20 → 3.00–6.10 3.00–3.20 covered by both
segments

5. Writing the Detailed Captions


A detailed caption describes the specific action in one segment. Unlike the high-level caption, detailed
captions do name the working hand — that level of precision is what the model learns from.

Rules
● Describe only what you see — no intentions, no guesses about goals. If you cannot see it, do not
write it.
● Name the working hand — say which hand performs the action. If a hand is resting, ignore it.
● Distinguish identical objects — when several similar objects are present, name the specific one as
simply as you can. If they are identical, use position: “The right hand picks up the far-right cup.ˮ If
they differ, use the distinguishing feature: “The right hand picks up the blue cup.ˮ
● Use simple present tense — write each action as it happens, with consistent structure and correct
grammar: “The right hand lifts the lid,ˮ not “lifted the lidˮ or “is lifting the lid.ˮ
● Reuse repeated captions — when an action repeats, reuse its caption from the Label Bank (Ctrl + B)
instead of retyping it

Examples
✓ “The right hand picks up the container and places it on the table.ˮ
✓ “The left hand opens the drawer and the right hand removes the tool.ˮ
✓ “The left hand places the bottle inside the orange box.ˮ

✗ “Trying to inspect the object.ˮ


Why: “trying toˮ and “inspectˮ assume intent.

✗ “Move object.ˮ
Why: too vague — which hand, and what object?
Idle and invalid labels
When no meaningful work is happening in a segment, apply one of these labels instead of describing
an action:

Label Meaning Use when


NU Not Useful No meaningful task or action:
a person is present but not
working, or no hands are
visible
DO Distracted Operator Attention is off-task: phone
use, talking, or smoking
ID Idle / Preparation Waiting with no interaction —
e.g., waiting for a machine
process to finish

6. Using work platformʼs Pre-Captioning (AI) Agent


● Correct captions only as needed — if a generated caption is correct, keep it, even if it is slightly
more detailed than you would write yourself. Rewriting wastes time.
● Edit only the wrong part — fix the incorrect word or detail; do not delete the whole caption and
retype it. Edit when the caption is wrong, the action or object is wrong, or a visible action is missing.
● Reuse for repeated actions — check the first 3–4 generated captions carefully. If they are right,
copy them to the matching segments. If they need the same small fix, correct once and reuse from
the bank.
● Use your judgment — the agent drafts; you verify. A confident-sounding caption can still be wrong.

The Label Bank populates as you write segments, so it will be empty at the very start of a video.

7. Common Shortcuts
Key Function
Shift + ↑ Set preview speed faster
Shift + ↓ Set preview speed slower
C Create a new segment per each action or
produce a detailed caption
G Fetch auto-generated captions
Shift + G Trigger caption generation if it didnʼt start
automatically
1/2/3 Select a model generated caption options
from the options keypads available
Ctrl + B Open or close the Label Bank
Ctrl + ← / → Jump to the first or last frame of the video
Alt + H Show or hide the caption overlay
F Toggle the Failure indicator
R Toggle the Recovery indicator
B Toggle the Bad Data indicator

More shortcuts are shown in the platform.


8. How Your Work Is Evaluated
Performance is measured on two things:
● Accuracy — the accuracy of your time-segmentations and the quality of your captions. Points are
deducted for untrue, inaccurate, or ambiguous captions, and for inaccurate segmentation.
● Target completion — the number of assigned tasks completed within the expected timeline. We
values efficient annotators.

Performance expectations
Maintain an overall accuracy score of 80% or above across both segmentation and captioning scores.
Annotators who consistently fail to meet quality standards and productivity expectations, even after
feedback and training sessions, may not be considered for future projects.

9. Cheating
Our work platform has a strong anti-cheating policy. Annotators have been caught attempting to cheat
and share answers in the past, and cheating attempts are continuously monitored using both proprietary
AI and a dedicated team.

Any annotator found cheating will be immediately removed from the project and permanently barred
from future work with Encord.

10. Final Checklist


High-level caption
☐ Simple English
☐ Clear task summary
☐ No hand or operator mention
☐ No technical jargon (No industrial terms needed, simple english works)

Segments
☐ Every change of action/event covered
☐ No gaps
☐ No overlaps

Detailed captions
☐ Action-focused, visible actions only
☐ Names the working hand (Most Important)
☐ Simple and specific
☐ Idle labels (NU / DO / ID) used when needed
☐ Simple present tense throughout
Productivity For Exam
☐ About 5 minutes per one minute of video
☐ Shortcut keys used
☐ Correct auto-generated captions reused
☐ Only incorrect parts edited — no full rewrites

FINAL EXAM HACKS


01. How Your Final Exam is Scored?
Your final score depends on two things:

• Quality – Accurate segmentation and captions.


• Speed – Completing each tasks in about 4–5 minutes.

02. High Level Caption is Very Important.


The AI generates segment captions based on the main caption you write.

• Spend enough time writing a clear and accurate main caption.


• AI suggestions are meant to save time, but they are not always correct. Always review and edit
them if needed.
• Follow the 8-Step Workflow every time.

If AI captions look completely wrong, it usually means:


• The main caption needs improvement, or
• The segmentation is incorrect (for example, two actions are combined into one segment).

03. Best Way to Prepare


The best way to pass is practice and consistency.
Every annotator should:
• Keep the 8-Step Workflow and Hotkeys Cheat Sheet nearby (printed or on a second screen).
• Watch the introduction video on work platform.
• Read the platform instructions on work platform.
• Review the annotation guidelines carefully. (This document)
• Practice 15-20 training tasks before attempting the exam.

04. You Donʼt Need Perfect Technical Names


You donʼt need to know the exact name of every object.
Just describe what you see clearly and accurately.
Examples:

Cup for a mug → Acceptable ✓


Glass for a mug → Incorrect ✗
Bowl for a mug → Incorrect ✗
Also, refer to the Golden Reference Examples to understand what good captions look like.

You might also like