Skip to content
AI Primer
πŸ€– ML/AI

eval-engineering

langchain-aiby langchain-ai2 months ago1.2k

Inspect an agent repository and optional traces, interview the user, write reviewed Task Specs, build and audit Harbor tasks, and bootstrap reusable project World Knowledge Skills. Use for agent evals, benchmark design, Task generation, controlled Environments, synthetic data, Verifiers, Harbor runs, calibration, or continuous benchmark maintenance.

Install

npx skills add https://github.com/langchain-ai/langchain-skills --skill eval-engineering
Show step-by-step
  1. 1

    Open your terminal

    • Mac: Press ⌘ Space, type "Terminal", press Enter
    • Windows: Press Win R, type "cmd", press Enter
  2. 2

    Paste the command above and press Enter

    Use the Copy command button, then paste in your terminal (Mac: ⌘V, Windows: Ctrl V).

  3. 3

    Restart Claude Code

    Close and reopen Claude Code, or start a new session, so it picks up the new skill.

Where it lives
~/.claude/skills/langchain-ai--langchain-skills--config--skills--eval-engineering/
β”œβ”€β”€ SKILL.md
└── ... (skill resource files)
View on GitHub

Comments

Always review skill code before installing. Third-party skills may contain scripts that run on your machine.

Related skills

πŸ’» Developer Tools
New

dynamic-workflow

Plan-in-code fan-outs, adversarial verification, waves.

by NousResearch Β· 6 days ago248k
πŸ€– ML/AI

comfyui

Generate images, video, and audio via diffusion workflows.

by NousResearch Β· 4 months ago248k
πŸ€– ML/AI

hyperframes

Render MP4/WebM videos from HTML compositions.

by NousResearch Β· 4 months ago248k
πŸ’» Developer Tools

claude-api

Reference for the Claude API / Anthropic SDK β€” model ids, pricing, params, streaming, tool use, MCP, agents, caching, token counting, model migration. TRIGGER β€” read BEFORE opening the target file; don't skip because it "looks like a one-liner" β€” whenever: the prompt names Claude/Anthropic in any form (Claude, Anthropic, Fable, Opus, Sonnet, Haiku, `anthropic`, `@anthropic-ai`, `claude-*`, `us.anthropic.*`, `[1m]`); the user asks about an LLM (pricing/model choice/limits/caching) β€” never answer from memory; OR the task is LLM-shaped with provider unstated (agent/MCP/tool-definition/multi-agent/RAG/LLM-judge/computer-use; generate/summarize/extract/classify/rewrite/converse over NL; debugging refusals/cutoffs/streaming/tool-calls/tokens). SKIP only when another provider is being worked on (overrides all triggers): OpenAI/GPT/Gemini/Llama/Mistral/Cohere/Ollama named in the query; OR `grep -rE 'openai|langchain_openai|google.generativeai|genai|mistralai|cohere|ollama'` over the project hits (run this grep FIRST if no provider named β€” don't Read the file).

by anthropics Β· 5 months ago177.6k