🤖 AI Tools
Label Studio Labeling Project Builder for LLM Assisted Annotation: Labeling Config XML with Choices, Labels, and TextArea Tags, Annotator Guidelines, Pre-annotation predictions JSON with model_version, Gold Task Review, and Export Checks
Set up a Label Studio project where an LLM drafts labels and people correct them: write the labeling config XML from your label schema, turn edge cases into annotator guidelines, convert model output into the predictions import format with from_name, to_name, type, and value, seed gold tasks to catch drift, and check the export before training on it.
0Reviews
Prompt
Act as an ML data operations lead who runs annotation projects in Label Studio, imports LLM drafted labels as predictions so annotators correct instead of start from blank, and has thrown away a week of labels because the config names did not match the import file. Inputs: - What is being labeled (text, chat transcripts, images, PDFs as images) and the data field names in each task: [DataSchema] - The label schema: classes or entity types, single or multi choice, required fields, free text fields: [LabelSchema] - Edge cases and disagreements seen so far, with real examples: [EdgeCases] - How the LLM drafts labels today (model, prompt, raw output sample): [LlmOutput] - Team size, Label Studio edition (Community or Enterprise), and review process: [TeamSetup] - Output format: [Format] Generate: 1. A labeling config XML built from DataSchema and LabelSchema: an object tag (Text, HyperText, or Image) whose value is $field, then control tags (Choices with choice="single" or "multiple", Labels for spans, TextArea for notes) with a unique name and the correct toName, required="true" where the schema says so, and hotkeys for the most used classes. 2. A name map table: every control tag name, its toName, its result type (choices, labels, textarea, rectanglelabels), and the value shape it expects. This table is the contract the import file must follow. 3. Annotator guidelines from EdgeCases: one definition per class, a "use this, not that" line for each confusable pair, a rule for when to pick an abstain class, and two short worked examples per hard case. Keep it short enough to paste into the project instructions. 4. A converter spec that turns LlmOutput into a Label Studio tasks file: each task with data matching DataSchema, a predictions list with model_version and an optional score, and result items with from_name, to_name, type, and value exactly as the name map says. For span labels include start, end, text, and labels, with offsets checked against the source text. 5. A small converter script in Python that reads the LLM output, validates every label against LabelSchema, drops or flags invalid ones instead of guessing, and writes the import JSON. 6. A gold task plan: a set of tasks with known answers mixed into the queue, how they are labeled so reviewers can find them, and what to do when an annotator or the LLM disagrees with gold. 7. Review and agreement: based on TeamSetup, either the Enterprise overlap and agreement settings or a Community workaround (duplicate a sample of tasks, export, and compute Cohen's kappa offline). 8. Export checks before training: JSON export, count of tasks with zero annotations, labels outside the schema, prediction accepted without change versus edited, and a note on keeping model_version so drafted labels can be audited later. Constraints: - Use only fields and classes given in DataSchema and LabelSchema. Never invent classes. - Mark any feature you cannot confirm for the stated edition as "check your Label Studio version". No em dashes.
Instructions
Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.
Generated Output
This image was generated using the prompt above.

Examples
Example Input
DataSchema: support chat transcripts, each task has fields "chat" (full transcript text) and "ticket_id"
LabelSchema: intent single choice: billing_dispute, cancel_request, bug_report, how_to, other; urgency single choice: low, high; free text field "why" only when intent is other
EdgeCases: customers who ask how to cancel are labeled how_to by some annotators and cancel_request by others; refund asks inside bug reports
LlmOutput: gpt-4o-mini returns {"ticket_id": "T-1042", "intent": "cancel request", "urgency": "High", "confidence": 0.71}
TeamSetup: Label Studio Community, 4 annotators, 1 reviewer, 2,000 chats
Format: XML, guideline text, converter, checksExample Output
1) LABELING CONFIG
<View>
<Text name="chat" value="$chat"/>
<Choices name="intent" toName="chat" choice="single" required="true" showInline="true">
<Choice value="billing_dispute" hotkey="1"/>
<Choice value="cancel_request" hotkey="2"/>
<Choice value="bug_report" hotkey="3"/>
<Choice value="how_to" hotkey="4"/>
<Choice value="other" hotkey="5"/>
</Choices>
<Choices name="urgency" toName="chat" choice="single" required="true">
<Choice value="low"/>
<Choice value="high"/>
</Choices>
<TextArea name="why" toName="chat" placeholder="Only if intent is other" maxSubmissions="1"/>
</View>
2) NAME MAP
| name | toName | type | value |
| intent | chat | choices | {"choices": ["cancel_request"]} |
| urgency | chat | choices | {"choices": ["high"]} |
| why | chat | textarea | {"text": ["..."]} |
3) GUIDELINES
- cancel_request: the customer wants the account or plan ended now, even if they phrase it as a question.
- how_to: the customer wants steps to do something themselves and has not said they want to leave.
- Use cancel_request, not how_to, for "how do I cancel" when the chat shows intent to stop paying.
- A refund ask inside a bug report stays bug_report; mention the refund in why only if intent is other.
- Abstain: pick other and fill why when no class fits after reading the full chat.
4) IMPORT TASK (one shown)
{"data": {"chat": "...", "ticket_id": "T-1042"},
"predictions": [{"model_version": "gpt-4o-mini-intent-v1", "score": 0.71,
"result": [
{"from_name": "intent", "to_name": "chat", "type": "choices", "value": {"choices": ["cancel_request"]}},
{"from_name": "urgency", "to_name": "chat", "type": "choices", "value": {"choices": ["high"]}}]}]}
5) CONVERTER
ALLOWED = {"intent": {"billing_dispute", "cancel_request", "bug_report", "how_to", "other"}, "urgency": {"low", "high"}}
def norm(v): return v.strip().lower().replace(" ", "_")
for row in llm_rows:
intent, urg = norm(row["intent"]), norm(row["urgency"])
if intent not in ALLOWED["intent"] or urg not in ALLOWED["urgency"]:
flagged.append(row["ticket_id"]); tasks.append({"data": chats[row["ticket_id"]]}); continue
tasks.append(build_task(row, intent, urg))
"cancel request" and "High" normalize to cancel_request and high. Anything else is imported with no prediction and listed in flagged.csv.
6) GOLD TASKS
- 60 chats with agreed answers, ticket_id prefixed G-, mixed in at random.
- Reviewer filters on G- weekly. If an annotator misses gold on cancel_request vs how_to, send them the guideline line, not a score.
7) AGREEMENT (Community)
Duplicate 200 chats as a second task set, assign to a different annotator, export both, and compute Cohen's kappa per field offline. Overlap settings are Enterprise only; check your Label Studio version.
8) EXPORT CHECKS
[ ] No task with zero annotations
[ ] No intent outside the five classes
[ ] why is empty unless intent is other
[ ] Count of predictions accepted unchanged vs edited, per model_version
[ ] Gold tasks removed from the training export