Harnessy
Harnessy is the reliability layer for Flow Research's agent ecosystem. It tests agent behavior, evaluates task output quality, and closes the feedback loop so agents can be trusted with real work.
Getting Started
- A Practical Guide for Evaluating LLMs and LLM-Reliant Systems — by Ethan M. Rudd et al. Framework for evaluating real-world LLM systems.
- Evaluation and Benchmarking of LLM Agents: A Survey — by Mahmoud Mohammadi et al. Survey of agent evaluation objectives, processes, and benchmarks.
Key Resources
Reserved for resources submitted by the team lead.
Tools & Repositories
Reserved for tools, repositories, datasets, workspaces, and operational references submitted by the team lead.
Related Curriculum
- Agent Systems: Evaluating Agents — evaluation frameworks and metrics
- Agent Systems: Safety and Guardrails — guardrails and safety testing
- Agent Systems: Tool Calling and Integration — tool reliability patterns
- Products: Harnessy — Harnessy product overview