Back to Discover

#openai batch api

1 prompt found

OpenAI Batch API Job Planner: JSONL Request Files with custom_id, Structured Output Schemas, Status Polling, and Result Joins for Bulk Classification
๐Ÿค– AI Tools

OpenAI Batch API Job Planner: JSONL Request Files with custom_id, Structured Output Schemas, Status Polling, and Result Joins for Bulk Classification

PpromptstudioยทOct 5, 2026
No rating

Plan a bulk LLM job on the OpenAI Batch API instead of looping over the regular endpoint: one JSONL line per record with a stable custom_id, a strict JSON schema for the answer, file upload with purpose batch, the batch create call with a 24h completion window, status polling, and a script that joins output and error files back to your source rows by custom_id and retries only the failures.

Act as an ML platform engineer who runs large offline classification and extraction jobs on the OpenAI Batch API and has debugged batches that failed validation or came back with results nobody could match to their source rows. Inputs: - The task in one sentence and the labels or fields each record should produce: [TaskSpec] - Source data shape: file type, row count, the primary key column, and two sample rows: [SourceRows] - Model and endpoint the team has chosen (for example /v1/chat/completions or /v1/responses): [ModelEndpoint] - The system instructions and any few shot examples, as they work today on single calls: [CurrentPrompt] - Language and runtime for the scripts (Python with the openai SDK, Node, plain curl): [Runtime] - Deadline and how results will be stored (CSV, database table, warehouse): [ResultStore] - Output format: [Format] Generate: 1. Job design: why Batch fits or does not fit the deadline in ResultStore, how to split SourceRows into batch files that stay under the current per batch request and file size limits (mark the exact limits VERIFY in the docs), and the rule that each input file targets one model and one endpoint. 2. A JSON schema for the answer derived from TaskSpec, with strict mode, required fields, enums for labels, and additionalProperties false. 3. One complete example JSONL line with custom_id built from the SourceRows primary key, method POST, url matching ModelEndpoint, and a body containing the model, CurrentPrompt, the record, and the schema in the response format field for that endpoint. 4. Script steps in Runtime: build the JSONL, validate every line parses and custom_id values are unique, upload the file with purpose batch, create the batch with completion_window 24h and metadata, and poll status through validating, in_progress, finalizing, and completed or failed, expired, cancelled. 5. Result join: download output_file_id and error_file_id, parse each line, match by custom_id because output order is not guaranteed, parse the schema answer, and write to ResultStore with columns for status, error code, and raw text on parse failure. 6. Retry plan: build a new JSONL from only failed or expired custom_ids, and a spot check of a small random sample by a human before trusting the labels. Constraints: - Do not state prices, discounts, rate limits, or turnaround times as facts. Tell the user where to check them. - Never put API keys in files or examples; read them from an environment variable. No em dashes.