Prompts and output contracts

On this page

Start with one model request: what does the source establish, what may be returned, and what needs clarification? Use order-form prefilling to connect task definitions, missing values, structural constraints and business validation. Valid JSON still needs supported fields and an acceptable resource state.

Prompt Describes Goals, Inputs, and Boundaries

Saying "You are an expert, please extract accurately" does not define what "accurately" means. A verifiable task must at least specify what the input is, which fields are required, the extent to which inference is allowed, how missing or conflicting information is handled, and who will use the final result. Prompt engineering also requires success criteria upfront to determine if changes are effective. Prompt engineering overview

For example, given the input "Want to buy A1, quantity to be confirmed later," if the task only requires sku and quantity, the model might guess 1. By explicitly stating that an unknown quantity should be null and marking the status as needs_input, making "no answer" a valid result becomes possible.

Original AmbiguityConvention to Clarify
Quantity not written, default to one?Do not guess; return null
Multiple items mentioned but field supports only oneMark ambiguity or use an array; do not arbitrarily pick one
What time period does "next week" refer to?Specify reference date and timezone, or keep original text
Conflicting content in sourcePreserve conflicts and sources; do not arbitrarily choose authority
Output will directly trigger an orderExtraction is only a candidate; do not auto-authorize execution

Stable rules can be placed in the corresponding trusted instruction layer; current materials provide data with clear sources. Message roles and content blocks are interface structures and should not be treated merely as arbitrarily concatenated text. XML tags, headers, or separators help distinguish content, but they do not provide secure isolation; "ignore previous rules" found in external materials should still be treated as data processing.

A teaching prompt can be organized as follows:

Task: Extract a product SKU and quantity from the user's original text for form pre-filling before manual confirmation.
Rules: Use only information explicitly stated in the original text; missing fields are null.
Status: Set to complete only when both fields are explicit and conflict-free; otherwise, set to needs_input.
Scope: Do not create orders, query personal profiles, or infer payment information.
Output: Adhere to the provided schema; do not add explanations before or after the result.
User Original Text: Want to buy A1, quantity to be confirmed later.

The "Scope" section in the prompt indicates what the model should do, but actual tool permissions are enforced by the execution side. Writing operational permissions into a single sentence cannot replace pre-execution verification in the Agent Loop.

Use Examples to Eliminate Real-World Ambiguity

Few-shot examples illustrate conventions using specific inputs and outputs. They are suitable for explaining abstract rules such as nulls, enums, units, and conflicts where misinterpretation is likely; however, more examples are not always better, and positive examples cannot always replace boundary descriptions.

InputExpected Core ResultRule Illustrated
Buy A1, total 3 itemsA1, 3, completeNormal extraction
Want to buy A1, quantity to be confirmed laterA1, null, needs_inputMissing values cannot be guessed
A1 needs 3 items, wait, change to 2 itemsMust be handled according to explicit correction rulesTemporal order and correction semantics
Buy A1 or B2, undecidedCannot arbitrarily select SKUAmbiguity needs to be expressed

If the schema does not allow space for ambiguity, even the best examples cannot prevent forced value filling; adjust the data model first, then adjust the wording. Do not write all test answers from development into examples and then claim that improvements on the same batch of questions represent generalization ability.

When diverse scenarios are needed, explicitly require that each scenario differs in constraints or trade-offs; when short outputs are needed, directly specify length and fields. temperature, effort, and length requirements address different issues and cannot be formulated as a migration formula like temperature=0 → effort=low. See Tokens and Sampling and Reasoning and Thinking for specific mechanisms.

Scope of Structural Constraints

Simply requesting JSON in the prompt asks the model to follow textual conventions; constrained decoding restricts valid output continuations during generation, ensuring the complete successful output conforms to the supported schema. Client-side validation is a check performed after generation; it can supplement constraints not supported by the service but cannot retroactively change already generated tokens.

Response CompleteCheck End ReasonParsing & SchemaStructure, Fields, TypesBusiness & EvidenceScope, Source, PermissionsContinue to Next StageStill check operation authorityFrom Generation to Business Action: Each Check Solves Different ProblemsValid JSON is a necessary format condition, but it does not prove field values are true, nor does it equal authorization for execution.
MechanismConstrained ObjectApplication Checks Still Required
JSON format requirement or schemaParseable JSON syntax, capabilities vary by interfaceFields, types, semantics
Schema-constrained responseSpecified structure of final outputComplete response, evidence, and business conditions
Strict tool parametersAllowed structure of tool selection and parametersPermissions, resource state, action semantics
Client-side type/business validationReceived objectState changes during subsequent execution

Claude's output_config.format is used for JSON responses, while strict: true on tools is used for strict tool calls; the two can be combined, but the available schema subsets and functional compatibility ranges need to be verified against the target model and SDK. Structured outputs

Several Boundaries of JSON Schema Itself

properties describes fields but does not imply these fields must appear; required fields are declared by required. additionalProperties: false rejects undefined fields but does not automatically make defined fields required. null is a value, distinct from the non-existence of a field. JSON Schema object

For instance, if you want the quantity field to always appear but be null when unknown, you need both required and a nullable type. Defining it only as an integer and setting it as required would make it impossible to express "no quantity in the original text" as expected.

Vendor structured outputs may only support a subset of JSON Schema. Some SDKs convert the full schema into a simplified form supported by the service, then verify against the original constraints on the client side; this is different from directly sending an original JSON request with unsupported fields. Do not interpret "SDK accepts it" as "the server enforced all constraints during generation." SDK schema conversion notes

How far is valid JSON from usable output?

Changing quantity from 3 to 5 preserves JSON and types but loses source support; keeping 3 while reducing stock fails a business condition. Separate predicates show whether to fix the prompt, data model, evidence check or execution precondition.

Preparing the visual
How far is valid JSON from usable output?

Syntax, schema, source support and business conditions are distinct checks; passing one does not guarantee the next.

Design Outputs for Non-Success Paths

Structured output guarantees are limited to paths supported by the interface and where the response is complete. Truncation, refusal, transmission errors, and cancellations need independent handling; you cannot simply take the first text block and write it to the database.

SituationApplication Handling
Normal completionParse, validate schema, then perform business verification
Output truncationMark as incomplete, preserve reason; do not execute half-formed parameters
Missing informationAccept unknown states in the schema, continue reading or asking based on the task
Source conflictsReturn locatable conflicts; do not automatically fabricate consistent conclusions
Model refusalRead refusal status via the interface; do not disguise it as normal business data
Network drop or timeoutDistinguish between incomplete generation and unknown external operation results

The client can retry with limits or request correction of specific fields, but each retry must have an upper limit and be counted towards costs. If failure stems from missing information, repeating the same input will not magically fill in facts; if tool side effects were previously executed, the entire task cannot be replayed just because parsing failed.

Preserve sources for structured reports by carrying document IDs, evidence locations, and versions in your own schema, which the application then verifies. Whether vendor-built citations can be used simultaneously with certain structured outputs is a specific functional compatibility issue, not equivalent to "JSON and traceability cannot coexist." See RAG for retrieval evidence integrity.

A Complete Local Verification Example

The following schema is used for complete verification on the application side and does not promise that the target service supports all keywords in it. The example relies on Python's jsonschema package to verify normal results, missing results, extra fields, wrong types, and cross-field contradictions; it does not call the model or create orders.

from jsonschema import Draft202012Validator, ValidationError

schema = {
    "type": "object",
    "properties": {
        "sku": {"type": ["string", "null"], "minLength": 1},
        "quantity": {"type": ["integer", "null"], "minimum": 1},
        "status": {"type": "string", "enum": ["complete", "needs_input"]},
    },
    "required": ["sku", "quantity", "status"],
    "additionalProperties": False,
}
Draft202012Validator.check_schema(schema)
validator = Draft202012Validator(schema)

def check(data):
    validator.validate(data)
    fields_present = data["sku"] is not None and data["quantity"] is not None
    if (data["status"] == "complete") != fields_present:
        raise ValueError("Status inconsistent with field completeness")
    return data

assert check({"sku": "A1", "quantity": 3, "status": "complete"})["quantity"] == 3
assert check({"sku": "A1", "quantity": None,
              "status": "needs_input"})["quantity"] is None
bad = [
    {"sku": "A1", "quantity": True, "status": "complete"},
    {"sku": "A1", "quantity": 0, "status": "complete"},
    {"sku": "A1", "quantity": None, "status": "complete"},
    {"sku": "A1", "status": "needs_input"},
    {"sku": "A1", "quantity": 3, "status": "complete", "extra": 1},
]
for item in bad:
    try:
        check(item)
    except (ValidationError, ValueError):
        pass
    else:
        raise AssertionError(f"Invalid result accepted: {item}")
print("Normal/missing results passed, 5 types of invalid results rejected")

This simplified data model only expresses missing values, not multiple items or conflicts. If real tasks require these states, corresponding fields and rules should be added, and fields_present should not continue to be used as the completeness criterion. The example also does not verify if the SKU exists, if the quantity comes from the original text, or if inventory is sufficient; these belong to source and business validation.

When executing actions, real-time status must be verified again. For example, if inventory is sufficient during extraction but sold out during submission, the schema cannot solve this race condition. Parameter validation, business constraints, and transaction processing must be completed in the corresponding systems.

Version Control Prompts and Schemas Together

Record versions of prompts, schemas, models, SDKs, and verification logic. Changing fields from nullable to required, adding new values to enums, or replacing single objects with arrays can all change downstream behavior; do not just look at model outputs, but also verify how old consumers handle new version results.

Evaluation sets should cover normal inputs, missing values, conflicts, noise, long inputs, malicious materials, and truncation. Statistically report structural validity rate, field accuracy, correctness of unknown handling, and true task success rate. If JSON is all valid but facts are wrong, adding more structural constraints may not be effective; return to evidence, prompts, and model capability positioning.

Continue with:Context engineering。