Agent Evaluation — Task Success Rate, Eval Harnesses, Benchmarks
How to evaluate AI agents: eval harnesses, task success rate, benchmarks like GAIA and Tau-Bench, LLM-as-a-judge — and how it differs from pure model quality.
Autonome KI-Systeme, die Aufgaben mehrschrittig bearbeiten.
How to evaluate AI agents: eval harnesses, task success rate, benchmarks like GAIA and Tau-Bench, LLM-as-a-judge — and how it differs from pure model quality.
Short- vs. long-term memory in AI agents: context window, external memory stores and vector DBs — how agents keep context across sessions.
How AI agents work: from a single tool call through MCP, structured outputs and LangGraph to the question of when multi-agent setups actually pay off.
How an LLM uses tools: define a tool as a schema, the model picks the function and arguments, the result returns to the chat — the basis of every agent.
How guardrails secure agents: input/output filters, policy checks, tool boundaries, and prompt-injection defense, in relation to security and HITL.