The open-source workspace where humans and AI agents work as one team.
-
Updated
Jul 26, 2026 - Python
The open-source workspace where humans and AI agents work as one team.
OpenCode++: a Coding Agent Reliability Harness for OpenCode, adding context, edit boundaries, command evidence, verification gates, impact analysis, and repair loops.OpenCode++:面向 OpenCode 的 AI 编程可靠性增强框架,为其增加上下文管理、编辑边界、命令证据、验证门禁、影响分析与修复闭环能力。
Compose coding agents into workflows you can trust
Map, inspect, and audit AI-agent markdown ecosystems (skills, agents, commands, hooks) as a graph.
Local-first AI that can actually use your computer.
Guardrails for the agent harness. Regulated data stays on your machine, and every prompt and tool call is scanned before it runs.
Agentic eval framework for Laravel: golden datasets, LLM-as-judge, adversarial harness, regression detection
Cyberful is an open-source application-security workbench for discovering, exploiting, verifying, and remediating vulnerabilities.
Source-available Harness AI / Agent OS for teams building governed agent applications - with Nexus orchestration, policy-controlled tools, RAG, memory, trace/evidence, HITL, and runtime boundaries.
An autonomous, self-learning AI coding agent that handles complex, multi-step development tasks - directly from your terminal, with the model of your choice.
Agentic Ops Harness | Superpowers Like AI Harness for Ops. Currently works with Hermes Agent
Shovs LLM OS is a local-first research runtime for studying agents that plan, use tools, remember, verify, and explain what happened.
OpenVelo, your new autonomous AI engineering team. OpenVelo is an open-source, fully automated software development orchestrator that transforms how you build software—from conversational ideation directly to tested, production-ready Pull Requests on your Linux-based infrastructure.
The rules layer: a Copier template for language-independent, contract-first agentic coding harnesses.
Pluggable DeepEval scaffold for RAG, agents, and LLM apps across Anthropic, Bedrock, Azure OpenAI, and Vertex. Ships traceability, test synthesis, safety/PII gating, multi-turn conversation eval, agentic tool-use scoring, JSON validation, judge benchmarks, hyperparameter sweeps, and pytest CI — one Makefile target per feature.
Drop-in TruLens evaluation harness for tool-calling LangGraph agents. Swap LLM providers (OpenAI, Anthropic via LiteLLM, Bedrock, Cortex, Gemini, Ollama) with a single env var. Ships with the RAG Triad plus Plan Quality, Plan Adherence, Execution Efficiency, and Logical Consistency metrics.
The Animus Project - Central Directory
Provider-agnostic RAG evaluation harness powered by RAGAS with pluggable LLM and embedding backends.
Smallest readable coding-agent harness that scores on benchmarks. ~970 lines, 59.6% on Terminal-Bench 2.0.
Building AI developer tools in Rust 🦀 | MCP servers • Agentic CI/CD • Code Intelligence • Hodei Platform
Add a description, image, and links to the harness-ai topic page so that developers can more easily learn about it.
To associate your repository with the harness-ai topic, visit your repo's landing page and select "manage topics."