Sherlock AI logoSherlock AI

Video Annotation Guide

A practical guide to creating video annotations.

This guide explains how to annotate self-recorded ground truth interview videos.

The purpose of annotation is simple: record what can be seen in the video, what was actually happening during the recording, whether the behavior was cheating or benign, and how certain you are about the annotation.

For each relevant video segment, annotate the following five fields:

  • Observable Behavior
  • Ground Truth Behavior
  • Label
  • Confidence
  • Description

1. Observable Behavior

What does it mean?

Observable Behavior describes what you can directly see or hear in the video.

Only describe what is actually visible or audible. Do not guess what the candidate is doing based on the recording scenario.

Ask yourself:

If I knew nothing about how this video was recorded, what could I confidently say just by watching and listening to it?

Good examples

  • The candidate repeatedly looks toward the right side of the screen.
  • The candidate looks down below the camera for several seconds before answering.
  • The candidate appears to scan text from left to right while looking away from the camera.
  • The candidate repeatedly looks toward the same off-screen area.
  • The candidate pauses for several seconds before beginning an answer.
  • The candidate holds a phone and looks at it.
  • Another person's voice can be heard while the candidate is silent.
  • The candidate repeatedly touches their ear while listening.

Avoid making assumptions

Do not write:

  • The candidate is reading ChatGPT.
  • The candidate is using a second monitor to cheat.
  • The candidate is receiving AI-generated answers.
  • The candidate is searching Google.

Unless the tool or activity itself is clearly visible, these statements go beyond what can be directly observed.

For example, if the candidate repeatedly looks toward the right side of the screen, write:

The candidate repeatedly looks toward the same area to the right of the screen.

Do not automatically write:

The candidate is reading ChatGPT.


2. Ground Truth Behavior

What does it mean?

Ground Truth Behavior describes what the candidate was actually doing during the recording.

These are self-recorded ground truth videos, so we often know what was happening even when the camera cannot directly show it.

Use the recording instructions, scenario, or your knowledge of how the video was created.

Ask yourself:

What was the candidate actually doing during this part of the recording?

Example: AI assistance

The recording scenario says the candidate should read an answer generated by ChatGPT on another screen.

Observable Behavior

The candidate repeatedly looks toward the right side of the screen and appears to read before answering.

Ground Truth Behavior

The candidate is reading an AI-generated answer from ChatGPT on a second screen.

The first describes what a viewer can see.

The second describes what we know actually happened.

Example: hidden phone

The candidate was instructed to use a phone below the desk to search for an answer.

Observable Behavior

The candidate repeatedly looks down below the camera and pauses before answering.

Ground Truth Behavior

The candidate is using a hidden phone to search for an answer.

Even if the phone itself cannot be seen, the Ground Truth Behavior can still describe the phone use because we know it happened during the recording.

Example: normal note-taking

The candidate was instructed to think normally and write notes on paper.

Observable Behavior

The candidate repeatedly looks down while writing.

Ground Truth Behavior

The candidate is taking handwritten notes and is not using any unauthorized assistance.

This type of example is especially important because it may look suspicious even though nothing improper is happening.

Do not guess

Ground Truth Behavior should come from known information about the recording.

If you do not know what the candidate was actually doing, do not infer it only from their appearance.

For example, repeatedly looking off-screen does not by itself prove that the candidate was using AI, Google, a phone, or another person.


3. Label

What does it mean?

Label describes whether the behavior was actually cheating, benign, or impossible to determine.

Use one of these three values:

  • Cheating
  • Benign
  • Unclear

The Label is based on what actually happened during the recording, not on how suspicious the behavior looks.


Cheating

Choose Cheating when the candidate is intentionally using unauthorized assistance as part of the recording scenario.

Examples:

  • Reading answers generated by ChatGPT.
  • Searching Google for an answer when external resources are not allowed.
  • Using a hidden phone to find answers.
  • Receiving answers or hints from another person.
  • Reading unauthorized notes or reference material.
  • Using another unauthorized application, device, or tool.

Example:

Observable Behavior

The candidate repeatedly looks toward the same off-screen area before answering.

Ground Truth Behavior

The candidate is reading answers generated by ChatGPT on a second screen.

Label

Cheating


Benign

Choose Benign when the candidate is not cheating.

This includes normal behaviors that may look suspicious.

These examples are particularly important because they help distinguish real cheating from normal interview behavior.

Examples:

Example 1 — Taking notes

Observable Behavior

The candidate repeatedly looks down toward the desk.

Ground Truth Behavior

The candidate is writing notes on paper.

Label

Benign

Example 2 — Reading the interview question

Observable Behavior

The candidate repeatedly looks toward the right side of the screen.

Ground Truth Behavior

The interview question is displayed on the right side of the candidate's monitor.

Label

Benign

Example 3 — Thinking

Observable Behavior

The candidate looks away from the screen and remains silent for several seconds.

Ground Truth Behavior

The candidate is thinking about the question without using any external assistance.

Label

Benign

Do not mark a segment as Cheating simply because the behavior looks unusual or suspicious.


Unclear

Choose Unclear when you genuinely cannot determine whether the segment should be labeled Cheating or Benign.

Examples:

  • The recording notes do not clearly indicate what the candidate was doing.
  • It is unclear exactly when the cheating behavior started or stopped.
  • The candidate was supposed to use an external resource, but you cannot confirm whether they had started using it during this segment.
  • Important information about the recording scenario is missing.

Do not guess just to avoid using Unclear.

If you are unsure about the ground truth, Unclear is the correct label.


4. Confidence

What does it mean?

Confidence describes how certain you are that your annotation is correct.

Use one of these values:

  • High
  • Medium
  • Low

Confidence does not describe how suspicious or serious the behavior is.

It only describes how certain you are about your annotation.


High

Choose High when you clearly know what happened.

Examples:

  • You recorded the video yourself and know exactly what the candidate was doing.
  • The recording scenario clearly states that the candidate should use ChatGPT during this section.
  • The candidate's behavior and the recording instructions clearly match.
  • You are confident about both the behavior and the selected time range.

For most well-controlled self-recorded ground truth videos, High should be the most common confidence level.


Medium

Choose Medium when you know the general behavior but are uncertain about some details.

Examples:

  • You know the candidate used ChatGPT during this section, but the exact moment they started reading is difficult to determine.
  • You know a phone was being used, but part of the behavior happened outside the camera frame.
  • The overall behavior is clear, but the exact start or end time is uncertain.

Low

Choose Low when important information is missing or you are not confident that the annotation is correct.

Examples:

  • The recording instructions are incomplete.
  • You cannot determine whether the candidate had already started using the external resource.
  • The video quality makes the behavior difficult to see.
  • You are unsure whether you selected the correct segment.

Do not increase Confidence simply because you need to complete the annotation.


5. Description

What does it mean?

Description provides a short explanation of the complete event.

It should connect what is visible in the video with what we know actually happened.

Unlike Observable Behavior, the Description can include both the visible behavior and the known ground truth.

Keep it short and clear. Usually one or two sentences are enough.

Ask yourself:

What happened during this segment, and what is important about it?

Cheating examples

The candidate repeatedly looks toward the same off-screen area and appears to read before answering. During this segment, the candidate is reading an AI-generated answer from ChatGPT on a second screen.

The candidate repeatedly looks down below the camera before responding. The candidate is using a hidden phone to search for answers.

The candidate pauses and appears to listen before giving an answer. Another person in the room is providing assistance.

Benign examples

The candidate repeatedly looks down, which could appear suspicious in the video. However, the candidate is only taking handwritten notes and is not using external assistance.

The candidate repeatedly looks toward the right side of the screen because the interview question is displayed in that area.

The candidate remains silent and looks away for several seconds while thinking about the problem. No external assistance is being used.

Keep descriptions factual

Avoid unnecessary conclusions such as:

This is extremely suspicious and the candidate is clearly trying to hide their cheating.

Instead write:

The candidate repeatedly looks below the camera while using a hidden phone to search for answers.

The goal is to describe what happened, not to judge the candidate's intentions beyond the known recording scenario.


Example 1: AI Answer Assistance

Observable Behavior

The candidate repeatedly looks toward the same area to the right of the screen and appears to read before answering.

Ground Truth Behavior

The candidate is reading an AI-generated answer from ChatGPT on a second screen.

Label

Cheating

Confidence

High

Description

The candidate repeatedly looks toward the same off-screen area and appears to read before answering. During this segment, the candidate is reading an AI-generated answer from ChatGPT.


Example 2: Hidden Phone

Observable Behavior

The candidate repeatedly looks down below the camera for several seconds before answering.

Ground Truth Behavior

The candidate is using a hidden phone below the desk to search for answers.

Label

Cheating

Confidence

High

Description

The candidate repeatedly looks down below the camera before responding. During this period, the candidate is using a hidden phone to search for answers.


Example 3: Normal Note-Taking

Observable Behavior

The candidate repeatedly looks down toward the desk while writing.

Ground Truth Behavior

The candidate is taking handwritten notes and is not using any unauthorized assistance.

Label

Benign

Confidence

High

Description

The candidate repeatedly looks down while writing notes. Although the downward gaze may look suspicious, the candidate is only taking handwritten notes.


Example 4: Reading the Interview Question

Observable Behavior

The candidate repeatedly looks toward the right side of the screen and appears to read.

Ground Truth Behavior

The interview question is displayed on the right side of the candidate's monitor.

Label

Benign

Confidence

High

Description

The candidate repeatedly looks toward the right side of the screen because the interview question is displayed there. No unauthorized external resource is being used.


Example 5: Uncertain Segment

Observable Behavior

The candidate repeatedly looks toward an off-screen area before answering.

Ground Truth Behavior

It is not known whether the candidate was viewing an external resource during this segment.

Label

Unclear

Confidence

Low

Description

The candidate repeatedly looks off-screen before answering, but the recording information does not clearly indicate what the candidate was looking at.


Key Rules to Remember

Observable Behavior = What can I directly see or hear?

Ground Truth Behavior = What do I know was actually happening?

Label = Was it actually Cheating, Benign, or Unclear?

Confidence = How certain am I that my annotation is correct?

Description = A short natural-language explanation connecting the visible behavior with the known ground truth.

The most important rule is:

Do not treat suspicious-looking behavior as cheating unless the recording ground truth confirms that cheating actually occurred.

For benign recordings, make sure to annotate behaviors that could realistically be mistaken for cheating, such as looking away, looking down, reading legitimate screen content, thinking for a long time, or taking notes. These examples are especially useful for distinguishing actual cheating from normal interview behavior.

On this page