Function Calling / Tool Use
How an LLM uses tools: define a tool as a schema, the model picks the function and arguments, the result returns to the chat — the basis of every agent.
Praktische Aspekte der LLM-API-Nutzung — Streaming, Caching, Rate Limits.
How an LLM uses tools: define a tool as a schema, the model picks the function and arguments, the result returns to the chat — the basis of every agent.
How prompt caching works with Claude, GPT-6 Astra and Gemini 3 Pro: write/read pricing, TTL variants, model differences, and when it actually pays off.
How Structured Outputs at Anthropic, OpenAI and Gemini enforce JSON schemas via constrained decoding — validated output for pipelines, not free text.
Practical API mechanics beyond pricing: streaming for UX, prompt caching against token cost, the Batch API for bulk jobs, rate limits without 429 drama.