0% found this document useful (1 vote)
1K views2 pages

Data Annotation Test

The document is a qualification test consisting of 40 questions divided into five sections: image classification, image bounding box, text classification, search quality & text correction, and guideline interpretation. Each section assesses the ability to accurately label images, draw bounding boxes, classify text sentiment, evaluate search result relevance, and interpret guidelines. The test includes various scenarios and examples to evaluate the annotator's skills in data annotation.

Uploaded by

Bayse Bacha
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (1 vote)
1K views2 pages

Data Annotation Test

The document is a qualification test consisting of 40 questions divided into five sections: image classification, image bounding box, text classification, search quality & text correction, and guideline interpretation. Each section assesses the ability to accurately label images, draw bounding boxes, classify text sentiment, evaluate search result relevance, and interpret guidelines. The test includes various scenarios and examples to evaluate the annotator's skills in data annotation.

Uploaded by

Bayse Bacha
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

FULL DATA ANNOTATION QUALIFICATION TEST (40 QUESTIONS)

SECTION 1 — IMAGE CLASSIFICATION (8 QUESTIONS)


Instructions: Choose the correct label (A. Cat, B. Dog, C. No Animal, D. Other Animal)
1. Image shows a black-and-white cat lying on a sofa.
2. Image shows a German shepherd looking at the camera.
3. Image shows a street with cars, no animals visible.
4. Image shows a horse running in a field.
5. Image shows a blurry picture; you think it's a dog but not sure.
6. Image shows a small kitten partially behind a curtain.
7. Image shows a stuffed toy dog.
8. Image shows a person holding a cat costume (empty inside).
SECTION 2 — IMAGE BOUNDING BOX (8 QUESTIONS)
Instruction: Draw a box around every real car. Ignore toy cars, reflections, drawings.
9. Parking lot with 5 real cars.
10. Poster showing a drawing of a car.
11. Street with one car and one motorcycle.
12. A car visible through a window.
13. A toy car on a table.
14. A real car reflected in a glass door.
15. Half of a real car is visible.
16. A sculpture shaped like a car.
SECTION 3 — TEXT CLASSIFICATION (8 QUESTIONS)
Labels: Positive, Negative, Neutral, Mixed
17. “I love the service. Everything was smooth.”
18. “I waited for two hours, but the result was okay.”
19. “Terrible experience, I’m never coming back.”
20. “It was fine, nothing special.”
21. “Good price but awful customer support.”
22. “Amazing! Fantastic! Perfect!”
23. “Not sure how to feel about this.”
24. “The product broke but they refunded me.”
SECTION 4 — SEARCH QUALITY & TEXT CORRECTION (8 QUESTIONS)
Part A: Search Relevance (Relevant / Partially Relevant / Not Relevant)
25. Query: “Best Ethiopian coffee brands”
Result: “History of coffee in Ethiopia”
26. Query: “How to download WhatsApp for Android”
Result: “WhatsApp Web login page”
27. Query: “Symptoms of vitamin D deficiency”
Result: “Vitamin D supplement stores near you”
28. Query: “Addis Ababa weather today”
Result: “Addis Ababa weekly forecast”
Part B: Text Correction
29. “She don’t like go there.”
30. “The informations is not accurate.”
31. “Him want to finish early.”
32. “This are the best choice.”
SECTION 5 — GUIDELINE INTERPRETATION (8 QUESTIONS)
Rule 1: Label a face only if real and clearly visible.
33. A real person smiling.
34. A face printed on a billboard.
35. A statue with a human face.
36. A real person wearing a mask.
Rule 2: Mark “Spam” if:
- Only emojis
- Gambling promotion
- No meaningful content
37. “■■■■■”
38. “Win $500 by clicking this link!”
39. “Hello, how are you?”
40. “?????????”

Common questions

Powered by AI

The guidelines direct annotators to correct errors by ensuring subject-verb agreement and proper use of singular and plural forms, such as changing 'She don’t like go there.' to 'She doesn't like going there.' Annotators must understand grammatical rules, including verb conjugation and pluralization, to make appropriate corrections that preserve the intended meaning while ensuring grammatical accuracy .

Guidelines instruct annotators to include only real objects in bounding boxes and distinguish these from fictional counterparts like reflections or drawings. Annotators need keen observational skills and an understanding of real-world context to discern subtle distinctions. Skills in visual analysis and an awareness of object characteristics are critical for effective annotation, ensuring accurate model training .

Items are included in a bounding box if they are real cars, not reflections, drawings, or toy cars. This criterion is important as it ensures only relevant and real objects are considered, which is vital for training machine learning models to recognize and work with real-world objects. Incorrect labeling could lead to models learning to recognize irrelevant objects or falsely identifying objects, reducing the accuracy and applicability of AI models .

Cultural and contextual understanding helps determine the intended nuances of search queries, such as 'Best Ethiopian coffee brands' versus 'History of coffee in Ethiopia'. Annotators must understand regional preferences and informational needs to evaluate relevance accurately. This cultural insight ensures search results meet user expectations and improve user satisfaction by delivering contextually appropriate information .

Labeling only real and clearly visible faces respects privacy by avoiding unnecessary identification of individuals who are not prominently featured. This approach helps mitigate potential misuse of personal data and aligns with ethical data annotation practices that prioritize individual privacy and consent, reducing the risk of identity-related exploitation or misuse .

The image classification guidelines require annotators to choose the most correct label based on visual content. Ambiguous cases, such as a blurry picture where a dog is suspected but not certain, present challenges as annotators must rely on the best guess or additional cues. Annotators face the difficulty of subjective interpretation, which can lead to inconsistencies in labeling, especially when images do not clearly fit into predefined categories such as 'Dog' or 'No Animal' .

Mixed sentiment content, such as a review stating both positive and negative aspects, presents a challenge as each sentiment must be considered. The guidelines label such cases as 'Mixed', accommodating the complexity of content that carries different emotional tones. The difficulty is ensuring that the sentiment analysis accurately reflects the overall text sentiment rather than focusing on isolated parts, which can lead to misrepresentation in sentiment analysis outputs .

Guidelines specify marking content as spam when it contains only emojis, gambling promotions, or no meaningful content to maintain information quality and reliability. This helps prevent users from being exposed to unfulfilling or harmful content that could degrade user experience and trust. Ensuring spam is filtered out maintains the integrity of communication platforms and aids in efficient content moderation .

Guidelines address indifference or uncertainty by labeling it as 'Neutral', avoiding bias towards positive or negative sentiment that might skew analysis. This labeling requires understanding nuanced language and recognizing when emotional intensity is absent. Accurate identification helps train models to handle ambiguous sentiment, ensuring outputs reflect genuine user intent without oversimplifying neutral statements .

Content is considered partially relevant when it closely relates but does not fully match the user’s search intent, as seen with 'Addis Ababa weather today' resulting in a weekly forecast. The implication for users is the potential shortfall in getting precisely targeted information, possibly leading to lower user satisfaction and inefficient information retrieval, as users may need to refine their queries for accurate results .

You might also like