The Agent Harness Hackathon: give AI models a license to act. August 24 to 30, 2026, with an NVIDIA DGX Spark and a Mac Mini among the prizes

The Agent Harness Hackathon

Give AI models a License to act

You get an agent working in an afternoon. Then you point it at something that matters and it can't reach your tools, can't run its own code safely, and can't be stopped before it does damage. Build one that can do all three, on TrueForge, TrueFoundry's open-source agent harness, and ship it the way you would at work: every change through a pull request, reviewed by Qodo before it merges.

Every submissionSubmit a project and you get a certificate of participation

Mission dossier

File TF-007

When
August 24–30, 2026. Monday 8 AM to Sunday 8 PM, London
Where
Take part ONLINE from anywhere, or join us in SF
Teams
Solo or up to 4 people
Prizes
$10,000 in prizes, including an NVIDIA DGX Spark and a Mac Mini, plus job interviews at TrueFoundry

Winners

Grand Winners

The best of everything handed in, online and in the room.

Best Use of TrueForge

NVIDIA DGX Spark

ONCALL

Elijah Umana

An on-call agent that takes an incident as far as a rollback, with every step of that rollback written down before it happens. The revert it expects is checkpointed before the push, a run interrupted halfway reconciles against the commit the remote actually has, and an incident survives a restart of the harness underneath it.

Best Code Quality

Mac Mini

CapyGuard

Nelson Lai

A guard that runs untrusted input in a sandbox and reports only what it watched happen there. Reading an artifact and ruling on it are split across separate agents, so the text under judgement cannot instruct the agent writing the judgement: a boundary the team found they had enforced by prompt alone, and rebuilt structurally, when their own review trail caught it.

Best UI

Apple iPads

QUEST

Julius Olsson

Best Use of Bright Data

AirPods 4

AgentRadar

Team AgentRadar · Biswajeet Sahoo, Priyanshu Solanki

An agent that treats a code review finding as a hypothesis rather than a verdict. It maps the finding onto the code graph, picks the tests that reach the line it names, runs them, and files the result as confirmed, unreproduced, uncovered, inconclusive or unlocatable.

Best blog post

Keychron Keyboard

Mayday

Anamika Singh

An incident responder that does the first twenty minutes of the work and none of the deciding. It pulls the metrics, reads the logs, checks what shipped and arrives with a diagnosis, the evidence behind it and the reasoning shown, then waits for a human before anything changes state. Recovery is judged on the latest raw telemetry rather than on an average the outage is still dragging down.

Online

Online Project Submissions

Handed in from anywhere between August 24 and 30, 2026.

286 projects

  • failsafe-trueforge

    Charannyan Kannan

    FailSafe gives an incident responder one narrow job: diagnose a checkout outage, ask for approval for the safest repair, and prove the system recovered. In the demo, an rc3 database-pool change pushes six checkout replicas toward a 200-connection Postgres ceiling.

  • Tech Titan

    Riju Paul

    Our project is an AI-powered Hand Matrix system that uses AI to analyze hand-related input and provide useful information or recommendations. It is designed to make the analysis process simple, fast, and user-friendly.

  • SchemaSentinel

    Mohit Pargaie

    SchemaSentinel is an agentic database migration safety system for developers and database engineers. It takes a proposed PostgreSQL migration, inspects the actual target schema, analyzes risks such as table rewrites and locking, validates the migration in an isolated PGlite…

  • Smart-E-Commerce-Support-Refund-Agent

    Shaik Mullasadik

    Smart E-Commerce Support & Refund Agent is an autonomous AI agent designed for online retail businesses to handle customer inquiries about missing packages and automate the refund process securely.In traditional e-commerce, customer support tokens for "order not received" take…

  • Scrap-verse-hackathon-

    Saniya Khan

    Focused on you data and codes

  • repomedic

    ANIRUDDHA ADAK

    RepoMedic is an autonomous open-source repository triage agent built on TrueForge. It scans repositories for real problems such as failing CI, broken README links, stale issues, and vulnerable dependencies.

  • SmartCart_AI-RAG

    DURGA PRASAD TATAPUD

    SmartCart is an AI-powered multilingual shopping assistant that helps users find and compare products using natural-language queries. It understands English, Telugu script, and Roman Telugu, extracts requirements such as budget, RAM, storage, processor, display, brand, and use…

  • trueforge-agent-assistant

    Dipayan Mukherjee

    SafeGuard Agent is an autonomous workspace assistant designed for developers and system administrators who need to securely audit directories, manage local files, and execute maintenance routines without risking security compromises.

  • AI-Incident-Responder

    Sachin Kumar

    The AI Incident Responder is an autonomous agent that handles production incidents end-to-end. When an alert fires (e.g., error rate spike on a service), the agent automatically investigates by querying service health metrics and deployment history, spawns parallel sub-agents to…

  • debrief-agent

    Yashaswini.V

    Debrief is a mission-dossier style web application that acts as a strict fail-safe for autonomous AI coding agents. Currently, allowing AI agents to navigate repositories is dangerous—they operate as black boxes that can blindly push broken code, leaving developers to clean up…

  • Poorva Shrivas

    Poorva Shrivas

    It is a safe agent on TrueForge that uses tools safely and can be stopped - solves unsafe agents problem for teams who need job-ready agents

  • KARAN YADAV

    Karan

    Projection Witness is a proof-carrying repair agent for teams operating event-sourced applications on PostgreSQL. It solves a subtle production failure where an out-of-order event is permanently skipped by a projector.

  • sentinelforge

    Aayush Sahu

    SentinelForge is a safety- AI incident responder designed for engineering teams. SentinelForge looks at repository incidents by using read- evidence. SentinelForge finds the root cause. SentinelForge creates a repair proposal that is limited in scope and gives it a fingerprint.

  • Aegis

    Bradley

    Aegis is an incident response platform for blockchain protocols. When something goes wrong on chain, like a treasury draining below the gas it needs to keep paying out, a contract acting strangely, or a big unexpected transfer, Aegis investigates on its own and comes back with a…

  • licence-to-patch

    Mayank Pande

    What it does: Licence to Patch is an agent that reviews your dependency-update pull requests and tells you which bumps you can trust — then, only with your approval, fixes the one you can't.

  • kernel-preflight

    Swapnil Dutta

    kernel-preflight is a GPU kernel-writing agent that cannot lie about the speedup — it submits kernel source and never submits a number. An agent that writes a kernel and reports its own speedup has both the incentive and the opportunity to be wrong in its own favour.

  • jarvis-ops-agent

    Sahil Mane

    Jarvis is a voice-first personal operations agent. You give it an outcome, not a sequence of app commands. It decides what evidence it needs across connected systems, does the work, and stops for human approval before anything irreversible.

  • divvy-forge

    Hitesh Pattanayak

    divvy-forge automates the research-and-draft step for dividend portfolio reviews. It reads my existing dividend holdings from a GitHub portfolio tracker, fetches live fundamentals (yield, payout ratio, FCF, EPS) from Screener.in and yfinance, runs two parallel subagents — one…

  • k8s-sentinel

    Anshul Bisht

    Autonomous Kubernetes incident triage agent built on TrueForge — real MCP tools, sandboxed analysis, subagent parallel dives, approval-gated remediation.

  • contain

    Deep Shah

    ContAIn is a secret-leak remediation agent. When a credential gets committed to a repository, someone has to work out three things by hand: does it still work, what can it reach, and will killing it break production.

  • casework

    Nick Sawinyh

    Casework works a transit data steward's feed-failure queue. A monitoring dashboard for a state's GTFS feeds shows dozens of failing feeds, but the person on rotation has a different question: how many problems is this really, and who do I write to?

  • arcadeops-mission-control

    Damien CREDOZ

    ArcadeOps Mission Control is a proof-carrying incident-response agent for operators who need AI to act without losing human control. It investigates a fictional degraded checkout service, delegates independent verification, validates a rollback in a Daytona sandbox, pauses for…

  • TrueForge

    Chirag Tankan

    AutoCompliance-Agent is an autonomous DevSecOps and regulatory compliance engineering agent designed to bridge the gap between compliance audits and code-level remediation.

  • FaultTrace

    Harshul Dwivedi

    FaultTrace is an autonomous forensic agent for vehicle diagnostics. Given one trouble code (e.g. P0171 "System Too Lean"), it runs a full investigation: gather evidence, enumerate competing hypotheses, compute a deterministic Bayesian differential, and spawn per-hypothesis…

  • checkout-services

    Sourjya Saha

    When production microservices crash in the middle of the night, traditional AI chatbots can only offer theoretical advice to tired on-call engineers. They cannot safely investigate, reproduce, or patch the live codebase without risking production outages.

  • Patchwork-AI

    Arpita Sharma

    Patchwork-AI is an autonomous developer agent designed to automatically detect code issues, simulate non-mutating patches, and streamline code reviews to maintain high software development standards.

  • auto-patch-agent

    Sai Rathod

    Auto-Patch Agent is an AI-powered incident response agent for production code failures. It takes an incident such as an HTTP 500 error and investigates the relevant repository code to identify the root cause.

  • DHARM PARESH KOSHIYA

    DHARM KOSHIYA

    SentryOps is an autonomous incident response agent built to monitor system alerts, inspect error traces via MCP tools, run automated regression tests in an isolated sandbox, and draft rollback patches.

  • Double-O-Harness

    Abhay Singh

    Double-O-Harness is an ultra-cost efficient, autonomous Open Source Intelligence (OSINT) agent capable of executing hundreds of concurrent web searches and compiling multi-page PDF dossiers with images.

  • Licence

    Leo

    Licence is a Harbor Pay on-call desk on TrueForge. Checkout error rate jumped from 0.4% to 8.1% after deploy 4c21 (PAGER-4419). The agent reaches a real MCP connector, fans out subagents across metrics, deploys, and logs, bisects the last four deploys in an isolate, and stops.

  • undox

    Manas Dutta

    Undox finds people-search sites that are leaking someone’s PII and helps drive opt-outs — but it never submits until a human allows the exact payload. It’s for anyone who cares about privacy, because removal today is slow, easy to mess up, and hard to undo once you hit Submit.

  • catherine (rin) pereira

    catherine (rin) pereira

    lifeagent is an approval-gated personal routine manager that reads your google calendar, sends push notifications for classes and events, reminds you to take medications at 8am and 8pm, nudges you toward sleep at 11pm, and runs a laundry countdown timer.

  • dossier

    Sanjeev Kumar S

    Dossier is an adversarial multi-agent idea intelligence platform that stress-tests startup concepts and technical proposals to find fatal flaws before capital or engineering effort is spent. Solves LLM sycophancy and ungrounded optimism.

  • evidenceforge

    cmdr-chara

    EvidenceForge is an evidence-gated control plane for investigating and repairing failed CI workflows. It binds every incident to an exact repository revision, coordinates three read-only diagnostic specialists, reproduces the failure in a Daytona sandbox, verifies the patch…

  • treasuryforge

    Romanch Roshan Singh

    TreasuryForge is an autonomous treasury agent that manages a simulated portfolio across cash, crypto (BTC/ETH), and NSE equities, built natively on TrueForge.

  • Ship_Safe

    Dhruvil Vora

    ShipSafe is an agentic PR test runner and code review assistant designed to make pull request validation safer and more automated. When given a pull request, ShipSafe analyzes the changes, coordinates agents to review the code, and runs relevant tests inside an isolated…

  • Sparkles

    Harieesh Ragaw SK

    ReleaseGuard is an AI-powered Release Safety Engineer built on the TrueForge agent harness. It autonomously analyzes software changes by inspecting pull requests and CI results through GitHub MCP, understanding repository architecture with DeepWiki, investigating production…

  • RoboOps

    Al Mahmud Samiul

    RoboOps is an autonomous AI Robotics Site Reliability Engineer (SRE) that diagnoses and resolves Autonomous Mobile Robot (AMR) fleet failures using TrueForge and MCP.

  • msdesk

    Anvit Devadiga

    Operations teams, engineering managers, and support leads waste hours every week manually pulling raw support metrics, calculating KPIs, and formatting status updates for Slack channels.

  • PriceMCP

    Tobias

    PriceMCP gives AI agents one MCP interface for trustworthy price comparison. It resolves an exact product variant before ranking normalized offers, then returns merchant identity, price, currency, freshness, availability, conditions, and provenance.

  • SentriX

    Anish Das

    SentriX AI solves the problem of finding and fixing security vulnerabilities in modern codebases. Traditional security tools often identify issues but leave developers to manually analyze and fix them.

  • Battery_Life_Prediction-Cost_Optimisation

    Anik Bera

    Modern battery management and manufacturing face three critical bottlenecks: The "Knee-Point" Degradation Trap: Batteries degrade non-linearly. Traditional battery management systems (BMS) rely on Coulomb counting or simple linear decay, causing them to miss the sudden…

  • leadforge

    Dharanidhara D J

    LeadForge finds leads and calls them to close deals. Find. Call. Sell without the huge manual work.

  • sentinelforge

    Navraj Singh

    SentinelForge is an autonomous Site Reliability Engineering (SRE) incident responder for on-call engineers and DevOps teams. When a production service breaks at 2:00 AM (like a checkout API returning 504 Gateway Timeouts), engineers usually have to wake up, manually comb through…

  • ForgeGuard

    Pranav Gupta

    ForgeGuard is a tool that checks whether an AI agent does tasks safely and correctly, not just whether it gets the right final answer. For example, if an agent is told to fix failing tests, it might actually fix the bug, while another agent could simply delete the tests.

  • Ctrl C + Ctrl V

    Harshit Saxena

    RefundGuard is an AI-powered refund investigation agent for customer support teams. It uses TrueForge and MCP tools to investigate customers, orders, payments, refund history, and policies, then recommends a refund decision.

  • TrueSRE

    Vishal Rajput

    TrueSRE is an enterprise-grade autonomous Site Reliability Engineering (SRE) multi-agent platform that detects, diagnoses, dry-runs, and remediates production Kubernetes outages in under 25 seconds—reducing Mean Time To Resolution (MTTR) from 45+ minutes.

  • self-correcting-integration-maintainer

    Keniel Maldonado

    Most code review catches what a rule already describes. The failures that actually ship are the ones nobody wrote a rule for. This is a Self-Correcting Integration Maintainer: an agent that inspects code it has never seen, is told nothing about what is wrong — the word "bug"…

  • GOOSEGOOSE

    Sami El-Figha

    The Gaggle is an adversarial multi-agent system for microbiome R&D. Instead of trusting a single AI model to produce a confident answer, it brings together independent AI scientists that argue for and against candidate hypotheses, retrieve live scientific evidence, run…

  • peel

    Srinivas Sivaratri

    Peel is a safety layer for sending .xlsx attachments from an owned mailbox. It preserves the exact attachment and SHA-256 hash, checks the workbook for hidden or unsupported content, runs verification in a fresh Daytona sandbox, and requires separate human approval before the…

  • Raghav Negi

    Raghav Negi

    ForgeOps is a terminal-native CLI agent for code review and incident debugging. It connects to a pre-configured TrueForge agent (forgeopsv1s) from any terminal - VS Code, GitHub Codespaces, Windows CMD, or Linux — and lets developers interact with their GitHub repositories using…

  • PMKVY

    SAI DUTTA ABHISHEK DASH

    SENTRY is an on-call incident responder that runs on TrueForge. When a production alert fires -- payment failures, 5xx spike, latency climb -- it investigates end-to-end over MCP: reads Prometheus metrics via Grafana, checks GitHub deploy history, correlates cause, and proposes…

  • forgeos-lite-trueforge

    Edson Dasilva

    ForgeOS Lite is a safety-first AI coding control plane for developers who want autonomous agents to build real software without giving them unrestricted authority over the original project. A user describes an idea in natural language.

  • sre-oncall

    Nitish Mane

    Two agents on the TrueForge harness. sre-oncall receives a Grafana alert, investigates over live MCP tools, correlates the failure to the ArgoCD sync and commit that caused it, and opens a revert PR carrying its evidence; a human merges, ArgoCD syncs, the alert clears, and it…

  • clinical-scrubber

    Biplab Bera

    Clinical Trial Data Scrubber & Analyst. Medical researchers spend hundreds of hours de-identifying patient data and running statistics before they can draft an FDA report.

  • mhcvi-live

    Shourya Sharan

    The problem no one talks about When disaster funding gets allocated, the algorithm picks winners by cost-per-person. Urban clinics get funded. Remote island communities get skipped. Not because they don't need help — because they're expensive to reach.

  • AI-Employ-Management-System

    Kajal

    Employee Management System is a web-based application that helps organizations manage employee information efficiently. It allows users to add, update, view, and delete employee records, making employee data management faster, easier, and more organized.

  • fixforge

    Ananya Unmesh Chaudhari

    FixForge is an autonomous software debugging agent. Given a bug report and a codebase, it investigates the actual project — listing and reading relevant files, reasoning about the root cause — then proposes a code fix.

  • AFTERMIND-WEBSITE

    MOHAMMED SAFIA TABASSUM

    What does your project do? AfterMind is an AI-powered emotional storytelling platform that transforms a user's personal thoughts, experiences, and emotions into meaningful visual stories.

  • clarity-agent-hackathon

    Ghulam Mustafa

    Clarity is an approval-first AI agent workspace that gives AI models a license to act safely under complete human control. It solves the risk of unchecked autonomous agents by introducing a Human-in-the-Loop safety layer.

  • airlock-mcp

    Himanshu Kumar

    Airlock audits what an MCP server actually does before an agent trusts what the server says about itself. It inventories declared tools, exercises each under a capped probe budget and compares observed behavior with published annotations.

  • verdict

    Himanshu Kumar

    Verdict turns intermittent GitHub bug reports into checkable reproduction records. Three bounded agents investigate the trigger, localize only what the evidence supports and design a regression plan.

  • Team Roferzzz

    Aishwary Gupta

    TrustForge is an autonomous AI Web3 auditor. It uses a TrueForge multi-agent system and a custom Python MCP server to automatically analyze Solidity smart contracts for vulnerabilities.

  • agent-change-gate

    Huang Chung Yi

    Agent Change Gate tests an agent instruction change before it reaches GitHub. It runs the current and candidate specs against the same 16 content-hashed scenarios five times, separates unstable results from regressions, then asks a person whether to create the branch, commit…

  • aegis-zero

    Aryan

    Aegis Zero is an autonomous DevSecOps agent that scans Python code for security vulnerabilities (SQL injection, hardcoded secrets, command injection), generates secure patches, and commits fixes to GitHub — but only after human approval.

  • true-trade

    Mayank Senani

    true-trade is a paper-trading agent for NSE stocks that runs on TrueForge. It screens a universe of 100 Indian large caps using free yfinance data, picks momentum candidates, and trades virtual money on a simulated clock — you start it on any past date, it trades on only the…

  • PaywallProof

    Vasu Bansal

    PaywallProof checks that a subscription provider's state matches the access a SaaS app actually grants. It runs four scenarios across API, browser, and application state: free access, paid access, scheduled cancellation, and post-expiry denial.

  • sentinental

    Mayank kumar upadhyay

    SENTINEL is an autonomous supply-chain security analyst built on the TrueForge agent harness. Point it at a repository and it takes dependency-vulnerability triage off your plate entirely — right up to the point where a human should decide. The problem. A CVE advisory lands.

  • secureops-guardian

    Jayesh Savaliya

    SecureOps Guardian is a human-controlled security incident-response agent for on-call, DevSecOps, and platform engineers. It connects a security alert to the exact GitHub change that caused it, checks the evidence, and prepares a safe, least-privilege fix.

  • SafeOps

    Sahil Warade

    SafeOps is an autonomous AI incident-response agent built on the core philosophy: "Investigate autonomously. Act only with permission." During critical production outages, on-call developers and SREs lose valuable time manually correlating metrics, deployment logs, and code…

  • dilate

    Devansh

    Dilate is an autonomous QA engineer for time-dependent bugs. It scans a codebase for logic involving dates, time zones, expirations, schedules, daylight saving time, and calendar arithmetic, then automatically generates targeted test scenarios and executes the application under…

  • APIx Predators

    Gautam Khosla

    Keyring is an access governance agent. You give it one instruction, such as "audit access for Ada Lovelace" or "offboard this contractor", and it fans out across your connected systems to inventory every grant a person or a machine still holds.

  • Parth Pinjarkar

    Parth Shrikant Pinjarkar

    AutoVault is an autonomous AI security operations agent that detects, investigates, and responds to ransomware attacks in real time — powered by TrueForge. The Problem It Solves Ransomware costs $20B+ yearly.

  • Orion Forge

    Arul Raj W

    Our UAV Disaster Management System is an AI-powered platform designed to assist emergency response teams during disasters. It analyzes UAV/drone footage using computer vision to detect people in affected areas and provides location-based information through an interactive…

  • sma-calibration-agent

    Juniad

    NEUTRINO identifies the seven parameters of a superelastic shape-memory-alloy (NiTi) constitutive model from a noisy tensile test. This is an inverse problem.

  • AgentHarness-Truefoundry-hackathon

    Damrukesh Daliparti

    An agent for a fictional CI/CD tool that answers support questions from a knowledge base, files GitHub issues, and reviews pull requests against documented conventions. Every write action (creating an issue, posting a review comment) pauses for human approval first.

  • codepulse-ai-agent

    Puneet bhardwaj

    CodePulse AI is an autonomous agent that fixes technical debt and broken test suites. It automatically scans repositories, isolates environments in a Daytona sandbox, applies patches using NVIDIA Nemotron, and opens verified GitHub Pull Requests for developers.

  • trueforge-incident-agent

    Bhagyashri Sinnarkar

    TrueForge Incident Response Agent helps engineers quickly investigate and resolve system incidents. It automatically collects logs, metrics, and incident data, analyzes the problem, and suggests possible solutions.

  • spyglass

    Vivek

    Spyglass is an evidence-engineering layer for AI incident-investigation agents: a Rust engine mines production logs, metrics, and deploy events into ranked, bounded, cited evidence bundles (served over MCP) instead of letting the agent page through raw telemetry then verifies…

  • trueforge-tally

    Vishal Jha

    It basically provides Indian SME's and MSME's to connect their accounting software like Tally directly to their agent. The navigation of tally is such a laborious task that it takes an CA 10 month just to master the software.

  • Circuit_Breaker_Agent_H

    Sandeep

    Circuit Breaker is a deterministic authorization control plane for AI-powered financial agents. It sits between an AI agent and the payment execution layer.

  • headcount

    Sarthak Agrawal

    HEADCOUNT is an idle game whose economy is human attention — and an AI agent designs the game while you approve its changes. Idle games model a company where labour is perfect: a machine makes 3.2 units a second forever, unsupervised.

  • Git Going

    Anugrah K S

    ARCHAEOLOGIST — Digital Forensic Workstation is an AI-powered agentic forensics and archaeology tool designed to reconstruct the context, architecture, and intent of abandoned or undocumented software projects.

  • PatchForge

    Praveen Rajan

    PatchForge is an agentic security platform that automates the end-to-end process of identifying, testing, and fixing software vulnerabilities (CVEs) within isolated, sandboxed environments.

  • trueforge-security-agent

    Darshan S

    TrueForge Security Agent is an autonomous security tool for enterprise data sanitization and forensic file recovery. It scans storage drives and soft-deleted recycle bin manifests, uses AI content inspection to classify sensitive data (SSNs, credit card numbers, API credentials…

  • Nutri-Ninja-5.0

    Jitendra Dewangan

    An AI-powered food scanner Android app. Scan any packaged food barcode, get instant health scores, ingredient warnings, personalized nutrition advice, and AI-powered diet chat — all stored locally on your phone.

  • Curious Pixels

    Tejas Krushna Rawool

    OpsForge is an AI-powered SRE command center that helps teams respond to production incidents faster. It automatically investigates logs, metrics, and deployments, suggests safe fixes, and verifies the results all while keeping humans in control of risky actions.

  • deadman

    Daniyal Ahmed

    DEADMAN is an AI SRE that doesn't just diagnose incidents it fixes production and undoes its own mistakes. It's for on-call engineers drowning in 3am alerts: it investigates from live signals, applies the fix, verifies it held, and auto-rolls-back if it didn't.

  • Crucible

    Bhagyesh Kamal

    The Crucible is a safety-first autonomous security validation agent. Given a controlled vulnerable target, it investigates, forms a vulnerability hypothesis, writes and tests its own PoC, and only after a human authorises the one sensitive step runs a controlled exploit to…

  • AEGIS-007-Autonomous-Incident-Commander-and-Harness

    Raj Tiwari

    Problem & Target Audience: A chatbot answers questions, but an autonomous SRE agent acts on production infrastructure: querying metrics, bisecting repositories, and executing rollbacks.

  • repo-debugger

    Shaik Nakeeb Mufeed

    Repo Debugger is an AI agent that helps developers fix bugs in GitHub repositories. Developers paste a repo URL and error description — the agent inspects the codebase using real GitHub tools, identifies the root cause with evidence from the actual code, proposes a specific fix…

  • bumpsmith

    Aryan Gorde

    bumpsmith migrates a Python codebase from pydantic v1 to v2, and it only keeps an edit that a green test run stands behind. The problem is narrower than "migration is hard".

  • Wemakedevs

    Golla Satvik Kumar

    Data Watchdog is an agent that checks a database for underperforming categories, judges whether the finding actually crosses a significance threshold, and only escalates to filing a GitHub issue when it does — every sensitive action, from running the SQL query to creating the…

  • RE-Agent

    Shubham Yadav

    RE Agent analyzes unknown Linux binaries safely — it runs static and dynamic analysis inside a hardened Docker sandbox, drafts a triage report, and pauses to ask for human approval before publishing anything to GitHub as a pull request.

  • auditforge

    Priyanshu

    AuditForge is an autonomous cybersecurity and DevOps incident triage agent that provides cryptographic audit trails for every decision it makes. When high-severity incidents strike (such as connection pool exhaustion, unmasked credential exposure, or CI/CD regressions)…

  • reroute-lg

    Ansuman Satapathy

    ReRoute-LG is an autonomous agent that triages supply-chain disruptions. When a typhoon or port strike threatens a shipping corridor, it checks live weather and news data, calculates how many days of inventory are left, evaluates alternate suppliers against cost, lead time, and…

  • trueforge-secure-agent-harness

    CHANDRA PAVANSAI

    The project, "The Safe AI File Guard: TrueForge Interception Harness," is an enterprise-grade incident response virtualization layer designed to allow autonomous AI agents a safe runtime environment.

  • handler

    Udit Jain

    HANDLER is the on-call engineer for someone else's training runs. A run dies at 03:14 and nobody notices until 09:00 — the GPU bills for six hours of producing nothing, and whoever owns the run starts their day reconstructing what happened from a log file.

  • Right2Erase

    Senali

    What does project do? Right2Erase is an AI-assisted system for handling GDPR “right to erasure” requests. It investigates a customer’s data across databases, file storage, billing systems, and logs, builds a deletion plan, and rehearses the plan in a disposable sandbox.

  • Sentinel

    Ronak Paul

    Sentinel is an AI SRE incident-response platform. When a production alert fires (via webhook from Grafana/PagerDuty) (We can create webhooks per project from the sentinel portal), it investigates root cause, verifies fixes in an isolated Daytona sandbox, and only ships changes…

  • deployguard-ai

    Anshu Gupta

    DeployGuard is an AI incident response agent for production and DevOps teams. It investigates service incidents by checking metrics, logs, and recent deployments, then identifies the likely root cause and recommends a fix.

  • SD007 - SOLO

    Sri Dharshan G D

    Odezzy AI audits and governs the MCP (Model Context Protocol) tools an AI agent can call. It scans every connected MCP server's tool definitions for prompt-injection patterns, leaked secrets, and schema mismatches; runs an LLM-based semantic check and embedding-based drift…

  • digital-declutter-agent

    Evangelin Blessy H K

    Digital Declutter Agent is an AI-powered filesystem assistant that investigates messy folders using MCP tools. It can recursively scan directories, identify duplicate files, read supported text and PDF files, process images through MCP ImageContent, summarize relevant documents…

  • codepulse-agent-ui

    Puneet

    CodePulse AI automates repository bug fixes in isolated Daytona sandboxes. It streams real-time terminal logs and enforces human-in-the-loop pull request approvals to safely accelerate software developer workflows.

  • hackathon

    Savar Bhasin

    I've built, The Squad, which is a control center for running a fleet of specialist agents. You give the squad lead a mission, it breaks those down into dependent tasks for agents such as software engineer, researchers, and work intelligence managers.

  • Proof of Chaos

    Mohit Kushwaha

    BRANCHPOINT is a safety and control layer for autonomous AI agents that need to take consequential real-world actions. Instead of letting an agent predict what will happen and immediately act, BRANCHPOINT makes it rehearse the action first.

  • portfolio-agent

    Zoya Muskaan

    It's an portfolio agent that turns your resume and GitHub username into an actual working portfolio site. Instead of just trusting whatever your resume says, it goes and checks your real GitHub repos first, writes the copy based on what it actually finds, builds a real site, and…

  • SentinelForge

    Akash

    SentinelForge is an AI SOC Incident Commander that autonomously investigates security alerts. It delegates work to Log, Network, and Malware investigator agents, gathers evidence through MCP tools, analyzes suspicious PowerShell artifacts safely, correlates findings, calculates…

  • bomb-squad-agent

    Anurag Thopate

    When production servers run out of memory or get stuck in cache deadlocks, SREs and backend teams need to fix it quickly. But giving an AI bot direct root access is dangerous because one wrong wildcard or bad command can delete active user sessions or important system files.

  • migration-sentinel

    Sumit S Chawla

    Migration Sentinel is an AI agent that investigates database migrations, dry-runs them in an isolated shadow sandbox, validates rollback, and gates production deployment behind human approval — evidence over predictions.

  • Fifty Shades of Agent

    Ashaya Sah

    For a Forex/Crypto trader, the project helps gather the technical data, technical analysis, news data and sentimental analysis all bound by Trueforge configured agent with agent harness and executing the trades if confirmed or instructed by a trader.

  • retrohost

    Sameer Goyal

    Published results are rarely re-run, so broken code goes unnoticed. Retrohost re-runs a paper's figures in isolated sandboxes via parallel subagents, classifying each REPRODUCED, PARTIAL or FAILED, then asking a human before publishing.

  • gatekeeper-trueforge-hackathon-

    vaishnavi adepu

    Gatekeeper is an approval-gated dependency upgrade agent for developers and small teams who know their dependencies are out of date but keep putting it off, because upgrading is tedious and you cannot tell what is actually safe without running something.

  • CampaignGuard

    Savanth JR

    My project "Campaign Guard" prevents broken ad funnels and unauthorized production overwrites by AI agents. When a marketing team wants to launch a new campaign, the agent autonomously verifies the destination URL health (preventing 404s) and generates sanitized short-links.

  • Jstreet

    Hemang Varshney

    MCP Vetter Agent is the security auditor of third-party MCP servers, running on TrueForge. It validates repos individually prior to agents connecting to them. The agent performs parallel probes (static AST analysis and dynamic).

  • Outsmart

    Harsh Singh

    Outsmart upgrades your npm dependencies and fixes what the upgrade breaks. Detection is a solved problem — Dependabot, npm audit and OSV all find the advisory. Repair isn't.

  • trueforge-hackathon

    Abhiraj kumar

    Anahita is a Hinglish-speaking, gothic-persona AI junior developer built on TrueForge. You give her a coding task by voice or text, and she works it the way a real junior dev would: she plans the change, writes and runs tests in TrueForge's sandbox, opens a PR, and waits for…

  • interviewos

    Priyam Srivastava

    InterviewOS is a research desk for one job: investigate a target software-engineering role, write a cited evidence ledger, compare it to a resume, and publish a four-file prep workspace (evidence.jsonl, brief.md, gaps.md, plan.md) only after a human approves.

  • Astella

    Naitik Morajkar

    Consensus Gap is an agent that takes any topic and researches two things in parallel: what experts actually say (from academic, institutional, and scientific sources) and what the public actually believes (from forums, social sentiment, and survey data).

  • AGENT

    Muhammed Nihal P n

    The Incident-Responder MCP Agent is designed to solve the critical problem of production outages that cost companies thousands of dollars every minute, serving as an essential tool for engineering and Site Reliability Engineering (SRE) teams.

  • passage

    Tushar Kumar Shahi

    Passage automates the health insurance pre-authorization process for hospital insurance coordinators in India. Right now, before a patient can get cashless treatment at a network hospital, a coordinator has to manually look up ICD-10 diagnosis codes, find the right procedure…

  • repo-rescue-agent

    Arthur Yousif

    Repo Rescue Agent helps developers diagnose and repair small repository issues. It inspects a repository, identifies a concrete bug, proposes a focused fix, runs verification, and produces a reviewable change for the developer.

  • migration-mate

    Sreeram Reddy Velagala

    MigrationMate is a safe, human-in-the-loop database migration agent built on TrueForge. Users describe database changes in plain English; the agent plans, tests in an isolated sandbox, asks for approval, opens a Qodo-reviewed GitHub PR with rollback scripts, merges, executes…

  • Edith-Hackathon-TrueForge

    Aarya Singh

    Our project is a multimodal AI desktop assistant designed to help users perform research, coding, productivity, and computer-based tasks from a single interface.

  • db-rag-agent

    Aman Kushwaha

    Deploy Detective (D2) is an approval-gated agent that mirrors an on-call engineer across two workflows: Incident response: Pulls real data via MCP, investigates deploys/alerts/metrics in parallel, bisects the culprit deploy in a sandbox, then pauses for human approval before…

  • AgentX

    Dinesh Jinjala

    AI agents with database access have a problem: the only thing stopping them from exporting customer emails is a system prompt asking them not to. Prompt injection, long contexts, or one bad completion and the data is out.

  • ARMAMENT

    Moazzam

    ARMAMENT is an on-call incident responder. When a service breaks, it does the investigation a human would do at 3 AM — reads container health, searches logs, checks memory and CPU, forms a diagnosis — and then stops.

  • breakglass

    Abhimanyu Goyal

    BREAKGLASS is an autonomous red-team security assessment agent that analyzes a target repository, discovers security-relevant indicators and attack surfaces, and generates actionable vulnerability hypotheses.

  • autofix-agent

    Priyanshu Singh

    AutoFix-Agent is an autonomous CI/CD failure auto-remediation and governance engine built for software engineering teams. Modern engineering teams lose 20% of their sprint velocity manually triaging broken CI/CD pipelines, deciphering runner stack traces, and writing regression…

  • student-performance-predictor

    Battena Anjali

    StudentAI is an AI-powered student performance prediction system that analyzes academic factors such as attendance, study hours, previous marks, assignment scores, internal marks, GPA, and backlogs to predict a student's expected performance.

  • Harmness

    Yash Gupta

    Portfolio Desk measures an investment portfolio against its target allocation, works out how far each holding has drifted, and proposes the trades that would bring it back into line.

  • 4Sum

    Chirag W

    DriftFix is an AI agent that turns breaking SDK changes into tested migration pull requests. In our demo, it upgrades a Python project from Stripe 14.3.0 to Stripe 15.6.0.

  • sentinelops-control-tower

    Ankit Anand

    SentinelOps Control Tower is an AI-native incident-response copilot. Given a security/operations alert (e.g. "payment failures spiking in checkout"), it: Investigates autonomously through read-only observability, deployment, and runbook tools (error rates, latency, logs, recent…

  • stores-triage

    Anmol Gupta

    Stores Triage is for a stores officer at a locomotive works. A spare part drops below its reorder level and an alert fires; he then has to answer a question the alert cannot — is this real?

  • Raphael-autonomous-engineer

    Satvik Mishra

    Raphael is an autonomous junior software engineer built on TrueForge, that reads a GitHub issue, implements the fix, validates it against tests, and waits for human approval before taking any irreversible action.

  • ColdCall

    Dhruv Mulay

    ColdCall is an incident commander for pharmaceutical cold-chain temperature excursions. When a shipment of medicine strays outside its labeled temperature range, someone must decide: release it, quarantine and retest it, or destroy it — an industry losing ~$35B a year to failed…

  • sentinel

    Jyotishman Pathak

    Sentinel is an autonomous Site Reliability Engineering (SRE) control plane and multi-agent incident management platform for distributed microservices. During production outages, on-call engineers are often overwhelmed by cascading telemetry noise and delayed root-cause analysis.

  • forty-two

    Pawandeep Singh

    Forty Two is an AI data analyst that lets users connect databases or upload CSV/Excel files, then ask questions in natural language. It autonomously discovers schemas, runs safe queries, analyzes data in a sandbox, and produces verified answers with durable tables and charts.

  • Outreach

    Pabitra Maity

    Outreach is an AI agent for cold outreach. You give it a company, role, and recipient - it researches the company live, drafts a personalized email, and creates a real Gmail draft once you approve it.

  • Dossier

    Bhargava Gumpula

    It applies to jobs for you and stops before sending anything. Its for anyone who is trying to apply for jobs. You name a company, and it reads their real job board and ranks every role against your resume, so a board with hundreds of openings becomes the handful you actually…

  • attest

    Kartik Shreekumar

    Attest checks whether an MCP server actually behaves the way it says it does. Every MCP tool can declare hints like readOnlyHint: true. Agents rely on those hints to decide what's safe to call without asking a human first. The problem is that nothing verifies them.

  • the-verifier

    Reet Singh

    The Verifier checks public claims by researching both supporting and conflicting evidence. It compares source dates, explains its conclusion, and requires approval before saving or exporting a dossier.

  • cloud-incident-responder

    Mrittiga M

    Cloud Incident Responder is an autonomous DevOps agent designed to help identify and respond to cloud/server incidents. It reads incident logs, analyzes the reported failure, identifies a likely root cause, and proposes a remediation action.

  • crucible

    shankar

    CRUCIBLE is an autonomous Site Reliability Engineering (SRE) and Incident Response platform that detects, diagnoses, sandbox-verifies, and remediates distributed microservice failure cascades.

  • BBBP-B3DB-Agent

    Biplab Ghosh

    NeuroPerm AI is an autonomous computational pharmacology and cheminformatics agent designed to screen Blood-Brain Barrier (BBB) permeability for Central Nervous System (CNS) drug discovery.

  • STRAW HATS

    Parth Gupta

    Sentinel is a prompt regression gate. When a PR changes a judging prompt, the TrueForge agent reads both versions through the GitHub MCP, runs a 10-case eval suite (subagents in the sandbox, scored with the sentinel-scoring skill), comments the comparison, and stops.

  • ByteBridge

    Simranjit Kaur

    CyberForge is an AI-powered SOC agent that investigates security incidents, correlates evidence, recommends containment actions, and requires human approval before taking action.

  • thinkfill

    Gopinath V

    ThinkFill is an AI-powered form filling assistant built for the TrueForge Agent Harness Hackathon. Filling forms repeatedly is a frustrating task, especially when users have to enter the same personal information again and again.

  • DataForge

    Mohith Tathineni

    DataForge is an AI-powered autonomous DataOps platform that detects data and pipeline problems, identifies their root causes, and helps automate their resolution.

  • insightforge-ai

    Ansh Dhanwani

    InsightForge AI is an autonomous business analytics agent that analyzes uploaded datasets to uncover trends, root causes, and actionable business insights.

  • Teen titan go

    Nil Lad

    Agent Guardian is an autonomous SRE incident-response agent that investigates production incidents by pulling real evidence — service metrics, error logs, and recent code changes — and proposes a root-cause fix, without ever being allowed to execute a destructive or unapproved…

  • Midnight

    Prathmesh Adsod

    Airstrong is airline disruption recovery software. It is built for situations where a problem that starts with one airport, aircraft, or flight quickly spreads across the rest of the operation.

  • Agent-Harness-Hackathon

    Manishika Agarwal

    FabGuard is an auditable maintenance-risk triage agent for semiconductor fabs and chemical plants. It reads structured maintenance events, calculates deterministic risk scores, generates reliability-focused recommendations, identifies evidence gaps and uncertainty, and prepares…

  • LocalRag

    Nizamuddien TI

    LocalRag is an AI career agent for job seekers. Instead of a chatbot that just answers questions about a CV, it actively works a job search: given a job description, it checks every requirement against the candidate's real CV — citing evidence for each skill rather than guessing…

  • applyguard-ai

    Sakshi Tripathi

    ApplyGuard is a human-in-the-loop AI agent that researches relevant career opportunities, identifies opportunities matching the user's interests, and presents them for review.

  • codegaurd

    Tuhin Banerjee

    A platform where you can check whether the codebase is safe for local installation or not just by putting the url.

  • DubSmash

    Agastya Khati

    ShutterFrame is a pre-production safety harness for database migrations. Today, teams often decide whether a migration is safe by reading SQL in a pull request.

  • MiloticX

    Peyton Li

    MiloticX takes a repo URL and walks through the README's setup steps one by one, running each command in an isolated sandbox to see what works and what doesn't. When a step fails it classifies the reason: a package got renamed, a flag was removed, or it's just a typo.

  • Byter

    MAYANK MAHAUR

    Byter is CI for bug reports. It turns incoming GitHub issues into deterministic proof, a tested candidate patch, and a maintainer-controlled draft pull request.

  • Endeavour

    Kush Yadav

    CHRONICLER is a persistent AI Dungeon Master for solo tabletop roleplay. Players speak or write a world into existence, choose an AI-generated quest thread, forge a character, roll D&D-style ability scores, and play through voice or natural-language actions.

  • resonance

    Kshitij Jain

    Resonance is for product managers who already have reviews and a dashboard that says “mostly positive.” Sentiment treats a polite 3-star “not bad” as okay; those customers often never file a ticket and then leave.

  • Ripple

    Manan Paliwal

    Ripple is a change-impact agent for supply chain and operations teams. When you need to rename a product SKU, Ripple reads your real operational database, runs deterministic impact math in a sandbox, shows you what is affected (purchase orders, shipments, customer orders…

  • tartarus

    Yerramsetty Sai Venkata Suchita

    1. Tartarus is an autonomous red-team and SecOps agent. It scans a GitHub repo, writes a real exploit, detonates it in an isolated sandbox to prove the bug is real, pauses for a human to approve, then opens a fix pull request. 2. Here the problem is scanners cry wolf.

  • aerosentinel

    Jyotishka chattopadhyay

    The project is primarily for organizations operating autonomous vehicles/robots where an AI can investigate incidents but should not be trusted to take dangerous actions without human oversight.

  • EraseGraph

    Sahil Rakhaiya

    EraseGraph is a privacy operations agent built around a simple idea: deleting personal data should end with evidence, not just a green “Done” message. When someone withdraws consent for a specific purpose such as using their data for AI model training—their information may…

  • kevinson

    Spencer Jireh Cebrian

    Cujo reviews pull requests by running them. A GitHub webhook starts one agent turn, the agent clones the PR into a throwaway Daytona sandbox, runs the test suite on base and head, writes and runs probes against the changed code, boots the app and hits it, and installs any newly…

  • Tracecause

    Vibushan Vinayak K

    TraceCause is an AI-powered causal investigation platform that helps users understand why a real-world outcome happened instead of just summarizing information. A user enters a company, event, or business outcome, and the system investigates it through multiple AI agents.

  • agent-harness-hackathon

    Ashutosh Maurya

    A retail store has 1,000 SKUs tracked in an Excel sheet. Instead of someone manually scanning it for low-stock items and calculating what to reorder, this agent does it — and pauses for human approval before writing anything.

  • Sudarshan

    Ujjwal Mohan

    Our project Connects 4 important tools to automate daily workflows safely, All with Human consent

  • querypilot

    Akshaj Manchanda

    QueryPilot is an evidence-first analytics and forecasting agent for revenue and operations teams at SaaS companies. It helps analysts investigate questions such as why qualified leads changed and what may happen next.

  • VSSP

    Sabarish SK

    OpsSentinel is an evidence-first incident response agent that investigates service failures autonomously but requires explicit human approval before taking destructive recovery actions.

  • pr-guard

    Sridhar

    PR Guard is a security-aware pull request review agent. You paste a GitHub PR URL. It reads the diff, runs the repo’s tests in an isolated sandbox, scans changed lines for security issues, drafts one review comment, then stops and waits for you before posting anything public.

  • ClanCode

    Gaurav Chaudhary

    ClanCode turns AI coding into a living game-like experience. Instead of watching an AI agent work inside a plain terminal, ClanCode represents your repository as a clan village.

  • cloudyyyy

    james gurung

    BlastShield is a safety layer between AI agents and production databases. It analyzes AI-generated SQL before execution, showing affected rows, cascade impacts, dependencies, and risk, then requires human approval before constrained execution.

  • portcullis

    Arpan Patra

    Portcullis audits an npm package before you let it into your repo, and it can't add it without your permission. The problem is that adding a dependency is the least careful thing most of us do.

  • PhoenixGen

    Riddhiman Roy Chowdhury

    Licence to Patch is an on-call agent that reads production errors from Sentry, investigates the code, reproduces the bug in a sandbox, creates a regression test, applies a minimal fix, runs the tests, and opens a pull request after human approval.

  • Ram And Co

    Surya

    AgentOpsML is an autonomous machine-learning experimentation system. It allows an AI agent to inspect a dataset, create reproducible train/validation/test splits, train multiple PyTorch model architectures, compare them using validation metrics, automatically select the best…

  • agent-harness

    Abhijit Mohanty

    OpenQuest is a human-in-the-loop agent for developers who want to contribute to open source but struggle to find the right repository, choose a suitable issue, or navigate the contribution process safely.

  • Orvexa

    Toufiq Farhan

    Orvexa is an autonomous database migration safety harness built for PostgreSQL engineering teams, platform engineers, and lead DBAs. The Problem: Modern CI/CD pipelines run migrations blindly: traditional runners consider a migration successful if PostgreSQL returns "exit code…

  • bountydesk

    Vaibhav Tomar

    BountyDesk is automated bug-bounty triage with a human sign-off. A vulnerability report arrives as a GitHub issue, an AI agent investigates it against a pinned target in an isolated sandbox, and the agent drafts a verdict with evidence: reproduced, not reproduced, or…

  • RECKON

    Hillary Ikhais

    RECKON is an agent built around one question: can I justify this action before I take it? It investigates real systems through MCP, builds an evidence-backed recovery plan, tests that recovery in a sandbox, red-teams its own reasoning, and then either asks a human to approve the…

  • ForgeCanary

    Dhiran

    ForgeCanary prevents MCP upgrades from silently breaking agent workflows. It replays successful jobs against both MCP versions, verifies the real external outcome, blocks regressions, approval-gates a scoped repair, reruns the jobs, and produces a release receipt.

  • UpgradePilot

    JAIVIGNESH GK

    UpgradePilot is a developer tool that verifies dependency upgrades before a pull request is created. It inspects a public GitHub npm/pnpm repository, shows outdated dependencies, runs the project’s real checks in an isolated sandbox, applies an upgrade, verifies it, and only…

  • codealongai

    Krishna Kartik Darsipudi

    Chat interfaces are linear, but learning, especially learning an unfamiliar codebase is not. Understanding code requires following relationships across functions, files, and concepts through multiple hops.

  • Spambots

    Janesh Kapoor

    SunoAI is a browser agent you drive in Hinglish that stops and asks before it ever writes anything. Ask it "analytics jobs dikhao" and it opens Chrome, navigates to the job listings, and reads the top results back to you, numbered so you can refer to one.

  • Managed

    Tyra KJ

    Airlink is an open-source remote control system for local AI coding agents running through True Forge. It lets developers start an agent on their workstation via a CLI or VS Code extension, pair it with a mobile phone using a six-digit PIN, send coding prompts remotely, view…

  • Warden

    Piyush Raj

    Warden is an incident response agent. When a production alert fires for a service, it investigates on its own: it pulls logs, metrics, and recent deploy history in parallel, then reproduces the suspected cause by actually running the candidate code in an isolated sandbox instead…

  • evidence-gated-runbook-executor

    Sahil Silare

    RunProof is an evidence-gated runbook executor for incident response. When an alert fires, its agent looks up the matching runbook, gathers only the evidence that runbook authorizes (logs, metrics, deploy history), runs a diagnostic script in a sandbox, assembles an evidence…

  • RailOps-Enterprise-DBMS

    Shambhavi M K

    RailOps Enterprise is a railway database management platform designed to centralize and simplify railway operational data. It provides an administrator-focused system for securely managing railway records through a structured, user-friendly interface.

  • omniforge

    Jay Joshi

    OmniForge is mission control for running autonomous agents against real infrastructure. It runs three specialized agents OpsForge for incident response, SecurForge for appsec, DataForge for data ops from one web cockpit.

  • ClipForge

    krish

    ClipForge is a TrueForge-powered video clipping agent. Upload a video, describe the moments you want in natural language, inspect and approve the agent's proposed tool call, and receive a rendered clip directly in the conversation.

  • sitrep-incident-commander

    Nagarjuna

    i have created SITREP, an incident responder agent. it analyses the server logs and other metrics by deploying sub agents. it suggest for change if it found any problem with the current deployment.

  • We-make-devs

    sepal sagar

    SafeOps AI is an approval-gated incident response agent designed for DevOps engineers and engineering teams. It autonomously investigates production incidents by analyzing monitoring metrics, application logs, recent deployments, and diagnostic results to identify the probable…

  • Null Pointers

    Ronak Gupta

    Autonomous, AST-aware, zero-downtime database migration & refactoring agent — built on TrueForge. SchemaForge treats a schema change as what it really is: a coordinated database and application-code change.

  • TypoBros

    Aditya Kumar Puri

    Researchers waste hours triaging papers: reading them, cross-checking claims, deciding what's trustworthy. Openwrite is an AI research agent that reads a paper for you and produces a Trail (what it did), a Coverage report, and a Claims↔Evidence table.

  • signalis

    Ramkumar R

    Signalis watches a company's CRM and website activity and tells a sales or marketing rep which leads are actually worth calling right now, and why. It reads in CRM rows and website events, figures out where each lead sits in the buying journey (early, mid, or late), explains…

  • trueforge

    Sakshi Andhale

    Trust Cop is an AI-powered security agent for MCP tools, built on TrueForge. It keeps a watchful eye on the tools your agents rely on tracking an approved baseline for each one and flagging anything that changes without warning.

  • AI_Repo_Analyzer

    Aayush Bansal

    AI Repo Analyzer is an autonomous multi-agent code auditing and technical due diligence engine. Manually vetting unfamiliar codebases takes 45–60 minutes per repo, while naive single-prompt LLMs suffer from high hallucination rates (34.5%) and missed vulnerabilities.

  • Agent-Harness-

    Rajan mishra

    Clearance is an agent for one job: a customer says they were charged twice. It looks up the order, proves the duplicate by running JavaScript in an isolated sandbox, then stops. process_refund is irreversible. Money does not move until a human stamps License.

  • aegis

    Harley Albert C. Buendia

    Aegis is an AI incident commander that investigates and fixes production incidents while keeping humans in control of high-impact actions. In our demo, a checkout service hits an 18.4% error rate.

  • flakebrake

    Dan Vellon

    FlakeBrake is a commitment firewall for humans and agents: it makes promise acceptance an admission decision instead of a model assertion, and it will not let an unsupported, stale, conflicting, duplicated, or unverified consequential claim cross an MCP boundary.

  • fuse

    Dhyaneesh DS

    Fuse is a governance and execution control plane for AI agents performing consequential work. Users visually author workflows that investigate an event, collect trusted evidence, rehearse an exact change, pause for human approval, execute through n8n, and independently verify…

  • We Crack Hacks

    Nitin Pathak

    VigilSRE is an autonomous AI Site Reliability Engineer (SRE) harness and incident response platform designed to eliminate production downtime. The Problem: When high-severity (P0) crashes occur (such as memory corruption, buffer overflows, or pool exhaustion), human engineers…

  • Barcelona

    Krishna.R

    Grid operators face incidents — storms, equipment failures, demand surges — where the "obvious" fix can cause a worse failure elsewhere on the network; GridMind's own hero scenario shows a proposed load transfer getting rejected in sandbox because it would have overheated a…

  • trueforge

    Mahijith Menon

    An agent that carries out GDPR "right to erasure" requests against a real production estate — and cannot destroy anything without a human saying yes.

  • 00-Flake

    Rohit

    00-Flake is an autonomous AI agent that detects, reproduces, diagnoses, and safely quarantines flaky automated CI tests so they stop blocking release trains. The Problem: Every software team experiences intermittent test failures in Playwright, Cypress, PyTest, or Jest.

  • GoatWhistle

    Mikhail

    Harness is a 24/7 autonomous paper-trading agent for Alpaca. It monitors market prices, company news, macro movements and portfolio risk, then proposes explainable trades. The main problem it solves is safely turning an LLM’s probabilistic reasoning into real financial actions.

  • Redline

    Rohan Jain

    What Does Your Project Do? Redline Guardian is a TrueForge agent that provides safe, auditable SQL change management for databases. It takes SQL statements from operators and runs them through a 4-role safety workflow: Ledger — Analyzes the SQL to identify affected tables, rows…

  • ai-dungeon-master

    ANUTTOMA SIL

    An AI Dungeon Master (AI DM) is an AI-powered game master that runs a role-playing game by acting as the narrator, rule engine, and world controller instead of a human DM.

  • trueforge_android

    Omkar Gothankar

    TrueForge Android Operator is a governed computer-use agent for a actual Android devices. Speak or type a task on the phone - "dismiss all my notifications except WhatsApp" and the agent operates the phone through its accessibility tree, recovering with a vision sub-agent when…

  • cleanroom

    Sreenath M Menon

    Every team has one spreadsheet nobody wants to touch. Three people have edited it. Dates are written three different ways. There are duplicate rows from a double import, money stored as text with dollar signs and commas, and totals that don't add up.

  • DrHiro

    Kresimir

    gathers all private health data in one place and allows user to control and act on it

  • ModelAudit-Agent

    Debabrata Pattnayak

    A scan report per model with confidence-scored findings, not a black-box pass/fail; a registry-wide view showing every model's current certification status automatic re-scan triggers whenever a new fine-tune or compressed variant of an existing model appears

  • agent-harness-submission

    Utkarsh tiwari

    My project is an AI-driven Data Pipeline Assistant designed to automate exploratory data analysis. It solves the problem of manual, repetitive data manipulation by writing and executing Python scripts (using Pandas and NumPy) directly within a secure environment.

  • Code Crafters

    Shashidhar Pawadashetti

    IncidentPilot is an autonomous incident-response and postmortem agent for backend services. When a production alert comes in (e.g. "connection pool exhaustion on order-service"), it fetches real diagnostic logs and metrics, diagnoses the failure, reproduces the crash in an…

  • vextorn

    Pranav Patil

    Autonomous AI robotics agents often succeed in idealized simulations but fail catastrophically in the real world due to hardware realities: sensor latency, packet drops, CPU thermal throttling (85°C), and actuator lag.

  • countersign

    Saharsh Tibrewala

    When an AI agent wants to delete something from your database, every harness just shows you the raw SQL and asks allow or deny. You can't actually tell what's about to happen. How many rows die, what cascades, whether you can undo it.

  • gatekeeper

    Anam khan

    Gatekeeper — The Trust Layer for Autonomous AI Agents Gatekeeper is an approval-aware autonomous agent that gives AI systems bounded autonomy when operating real-world tools.

  • Avengers

    Sri Nidhi Kuchana

    Issue Resolver Agent automates the manual loop of resolving a GitHub issue: reading it, digging through the codebase to find relevant files, diagnosing the bug, writing a fix, and opening a PR.

  • Runtime Rebels

    Eva Chen

    RootCheck is an AI agent security testing harness that evaluates how agents behave under adversarial conditions. It runs controlled security scenarios, such as indirect prompt injection, observes the target agent’s actual tool calls, and determines whether the agent followed…

  • drawn

    Mintu Gogoi

    An agent connected to an MCP server normally answers with a wall of JSON or a paragraph describing what it found. drawn makes it answer with interface instead. Ask for flights and you get fares you can click.

  • exitramp

    Siddharth Choure

    ExitRamp is a pre-migration gate for engineering teams replacing the model behind a tool-using AI agent. A final-answer-only evaluation can miss a dangerous regression: a cheaper model may sound correct while skipping required tools, using the wrong arguments, or making claims…

  • BlackBox

    Aayushi Gupta

    Sentinel is an autonomous crisis intelligence and emergency dispatch platform built for 911 dispatchers, emergency operations centers (EOCs), and first response commanders who face severe sensory overload and coordination delays during chaotic disasters.

  • rocky-code

    Aditya Sasidhar

    I built Rocky, a terminal coding agent for developers who want the speed of AI coding tools without giving them unrestricted access to a real working tree. Rocky delegates work to Codex, Claude Code, or OpenCode inside disposable containers.

  • ScholarAgent-Autonomous-Research-Reproduction-Agent

    Tanya Garg

    ScholarAgent is an autonomous AI research reproduction agent that helps researchers, students, and ML engineers verify whether the experimental results reported in a research paper can actually be reproduced.

  • trueforge-ai-buyer

    Tejas M

    Will mail it soon

  • trueforge-rote-fix

    Albert Joshua

    This is a general deep research agent that can deep search the web for content using sub agent spawning and produce analysis reports

  • scope-city

    Rajdeep Singh

    Give an AI agent a Stripe key so it can refund one charge, and it can refund all of them. The same is true of GitHub, Salesforce, or your database: every integration hands over a credential that can do far more than the job needs, and afterwards nobody can say what the agent was…

  • The Farmers

    Shanu Kumawat

    TrueStrike is an autonomous pentest agent that asks permission. Point it at a target you own and it recons the whole attack surface, proves every finding with a working exploit inside an isolated cloud sandbox, then stops and waits for you before anything destructive.

  • WeMakeRevenue

    Rudra Veer Singh Rathore

    Recoup is an SRE-style investigation agent for revenue, built on TrueForge. Point it at a failed-payment spike, a refund wave, or a renewal-risk signal, and it works the incident the way an SRE works an outage: pull evidence from Stripe, Sentry, GitHub, and tenant data via…

  • greenlight

    Aryan Kumar

    Greenlight is a CapCut-style cloud video editor with an AI Producer beside the timeline. Creators can drag, trim, split, mute, add transitions, mix audio, and undo normally, or ask the Producer to research, script, find licensed media, edit, render, and stage to YouTube.

  • vendor-scout

    Shivam Arora

    Vendor Scout is an autonomous procurement agent for hardware teams. You give it a part you need, your current supplier, quantity, price target, lead-time requirements, and technical constraints.

  • PodNot

    Akshat Arya

    ReDessIo is an autonomous spatial interior architecture and procurement agent harness built on the philosophy: "Watch it design. Approve before it spends." Problem It Solves: Generative image AI hallucinates beautiful interior concepts that can neither be purchased nor fit real…

  • sentinel-008

    Dhruv patel

    Sentinel·008 is an approval-gated incident response agent.When a payment-failure alert fires, it investigates the incident the way a senior SRE would: it reads dashboards, queries the database, searches logs, examines recent deploys, and runs diagnostic code inside an isolated…

  • operationsos

    Shivam Srivastava

    OperationsOS is an Business Operations Management Solution for a company run by AI agents a walkable 1-bit office where user is the CEO. Five specialists sit at desks doing real operational work on a live business: a Data Analyst querying the customer database, Market Research…

  • Blast Radius

    Nishad Mulay

    Blast Radius is a safe AI agent for deleting a customer’s personal data from a database. It first finds every record connected to that customer, such as addresses, uploads, support tickets, orders, audit logs, and order items.

  • Shaken, Not Staged

    FNU Smriti

    Production incidents demand fast decisions, but rushed remediation can worsen an outage. On-call engineers often need to manually gather metrics, logs, traces, deployment history, and ticket context before they can determine whether a rollback is safe.

  • agent-harness-automation

    Payas Dhone

    An AI agent system running on TrueForge that fetches real-time repository trends, executes analytical Python code inside an isolated Daytona sandbox, and enforces human-in-the-loop approvals before file operations

  • Warden

    Naman Jain

    Warden is an autonomous release gatekeeper built on TrueForge. It reviews GitHub pull requests by fetching real PR data through GitHub MCP, delegating security, style, and testing analysis to specialist subagents, running the PR’s exact commit inside an isolated Daytona sandbox…

  • aegis

    Abhishek Patel

    AEGIS is an autonomous runtime reliability agent for Node.js. It finds event-loop starvation (a CPU-bound task — massive JSON parsing, regex evaluation, sync crypto — blocking the single thread), reproduces the exact failure under a controlled load experiment, applies a targeted…

  • stamp

    Fawuzan

    Vendor reminders can look identical to new invoices, creating a real risk of paying twice for the same service. Using TrueForge, we built an AI agent that checks the books before drafting a payment-confirmation message.

  • Polaris-CLI

    Dilip Kumar R

    Polaris AI runs an ensemble of agents to coded implementations any Machine Learning paper. It aims to bring the knowledge of using coding agents to ML papers, to cut down on time taken for the implementations but instead focussing on running the experiements on our own platform…

  • docxy

    Arindam Majumder

    Docxy is an automated documentation and changelog pipeline powered by five specialized agents and human approval. Every push to main triggers agents that analyze code changes, map their impact on docs, update the docs, write changelog entries, and coordinate the final result.

  • trading-desk-sentinel

    N S Siddarth

    Sentinel Trading Desk is an AI-powered incident response and risk-monitoring system for paper-trading accounts. It connects to Alpaca through an MCP server and uses an AI agent to investigate unusual portfolio events by checking account state, portfolio history, recent orders…

  • Innovators

    Kamal

    SafeRun is a database guardian agent. In July 2025 an AI coding agent deleted a production database, ignored an instruction to stop, and claimed the data was unrecoverable. SafeRun is the layer that was missing.

  • picto

    Sahil Gupta

    I have noticed a problem of overwork of maintainers of open-source organisations. Due to AI development they are facing many sloppy prs/issues.

  • colophon-agent-harness

    Manann arora

    Colophon is a spec-first product video-generation harness with taste. An agent writes a video plan, Colophon enforces a closed motion vocabulary and runs timing, text overflow, brand-colour compliance, source-backed claims, and more.

  • fuseops-trueforge

    amit mishra

    FuseOps is an approval-gated incident-response agent for SRE and platform teams. When checkout payment failures spike, it gathers incident context, service health, deployment history, sanitized logs, and the runbook through an owned MCP control plane.

  • YukClara

    P Koti Darshan

    SENTINEL is an autonomous incident-response agent for Kubernetes. When a Prometheus alert fires (CrashLoopBackOff, API error spike, DB timeout), it loads the matching runbook, sends four sub-agents to investigate in parallel, the K8s state, logs, API metrics, and DB metrics over…

  • prism tech

    Dheeraj kumar

    When a bug/error shows up in production, this agent investigates it on its own, finds the root cause, proposes a fix, and verifies it in a sandbox — all without a human manually digging through logs and code, needing only one approval at the end

  • falcon-harness

    Saksham Mishra

    Falcon is a diff-scoped exploitation agent. AI coding agents ship new endpoints fast, and every so often one forgets an access-control check — and traditional scanners only guess a severity.

  • memoforge

    Aditya Gupta

    An analyst at a boutique venture advisory or VC firm receives a founder's pitch — a deck, an email, a set of notes — and has to turn it into a clean, consistent, house-format Investment Committee (IC) one-pager that a partner will read, trust, and act on.

  • clinical-ops

    Sakeena Patel

    This project is to automate healthcare administration reviews.

  • vetit

    Victory Lucky

    MCP servers are becoming more popular, in the same way it is being met with security issues and vulnerabilities, and according to AgentSeal 2026 report, 66% of 1,808 MCP servers had a security finding, that's roughly 1,193 compromised servers.

  • NOTEAM

    Arav Latiyan

    An agent that takes a suspicious email, detonates its links in a sandbox, and runs three parallel investigations (infrastructure, identity, history) into a verdict.

  • migration-harness

    Yashkumar Kavaiya

    Autonomously modernizes a bounded .NET service to Rust, then proves behavior before anyone can cut over. Compilation is not enough. Teams ship migrations that compile and still break production. This is for teams who need a cutover they can defend.

  • Null Garden

    Bhavay Goel

    The problem Webhook systems can tell you that a delivery failed. They cannot necessarily tell you whether the business operation failed. The provider records a failed delivery, but the business mutation has already occurred. Blind replay can perform it again.

  • forge-guard

    Shoumik Chandra

    ForgeGuard is an enterprise-grade AppSec governance layer designed for security engineers and platform teams deploying AI agents. It solves a critical vulnerability known as the "Lethal Trifecta" when an agent possesses simultaneous access to untrusted input, private data, and…

  • debugforge

    Priyanshu

    DebugForge is an autonomous debugging and code-repair system for developers. It tackles the growing difficulty of debugging complex and AI-generated code by reproducing failures, tracing likely root causes, generating surgical fixes, and verifying those fixes before applying…

  • graft

    Md Abid Hussain

    Graft is an agent memory for hackathon stacks. Every WeMakeDevs hackathon runs on a different stack, and the context an agent builds reading its docs dies with the session — so every event starts with re-reading the same documentation.

  • quartermaster

    Manu Mishra

    Coding agents claim. They say "I've fixed the failing test" and you find out in CI, twenty minutes later, that nothing ever ran. The claim is cheap because nothing forced the agent to execute anything. During development this agent did exactly that.

  • LUPIN-007

    Ramanpreet singh

    LUPIN is an agentic SRE assistant that replaces panic-driven guesswork with safe, reproducible diagnosis. The problem: When an alert fires at 3am, an SRE either blindly SSHs into production (risking making it worse) or sits paralyzed.

  • agentguard

    Jaival Suthar

    AgentGuard is a runtime trust layer for AI agents. I built it to control what an agent is allowed to do and independently verify what actually happened during execution.

  • taro

    Chetan Patil

    Today's AI agents talk to one person and work for one person. Real work isn't like that , what if we want to talk to the different people to schedule , coordinate to make a conclusion of a job , for example a single roofing repair means coordinating a homeowner, a project…

  • confirm-deny

    Anik Das

    CONFIRM/DENY turns an unverified bug report into a verified case file, or a defensible “cannot reproduce” with proof. Give it a GitHub issue URL. It runs the reported code in a credential-free sandbox, proves what ran, bisects to the first bad version, builds a structured case…

  • approvedeck

    Aarav

    Agent harnesses pause before anything irreversible, and that pause is the whole safety story. Today it is buried in a chat transcript. Run five agents and the approval that matters is three tabs deep and forty messages up, so people either miss it or rubber-stamp it later…

  • ROOK

    Jay Rietzke

    ROOK is an evidence-first AI incident commander built on TrueForge. It investigates incidents using least-privilege, read-only tools, correlates TrueForge tool calls and responses into retained evidence, and only promotes operational state when that evidence satisfies a strict…

  • rag-doctor

    Kaushal Francis J

    Checks current RAG system health and if rag is bad, runs diagnosis, and fixes.

  • repoguard

    1

    RepoGuard is an autonomous health and repair repository agent that inspects repository code and CI results and identifies the root cause of failing test cases, diagnose, make minimal repair and requires human approval before any repository change.

  • revenue-sre

    Siddharth Jha

    Revenue SRE is an AI-powered payment incident commander for online merchants using Razorpay. Payment dashboards show individual failures, but merchants can struggle to recognize when many failures represent one larger incident.

  • trueforge-proofboard

    Manuele

    Proof Board is a control plane (in Kanban style) for autonomous software delivery. It lets AI agents implement changes, independently reviews the actual sandbox artifact, and requires explicit human approval before publishing a pull request.

  • mohenjodaro

    Aryan Singh

    It helps in building testable evals for any robot on environements based on requirements and run eval sets in isolated sandboxes and relay the results in a research scoped UI desined for robotics engineers and researchers.

  • ResearchForge-AI

    Girivasan B

    ResearchForge AI is a multi-agent autonomous research platform that transforms a single query into a verified, citation-backed intelligence report. The system uses specialized AI agents, MCP tools, and sandbox execution to gather, analyze, validate, and synthesize information…

  • GuardForge

    Aman Singh

    Problem: AI agents given access to codebases have no safety layer — they can merge broken code, run untested fixes, or deploy to production without human oversight.

  • PATCHPILOT

    Swayam Prakash Sahu

    PatchPilot is an AI production engineer, built on the TrueForge agent harness — it automates the tedious middle of incident response while keeping a human in charge of the one irreversible step.

  • hello-agent

    Divyansh

    hello-agent (Repo Maintainer) is an agent that scans a GitHub repository for real, verifiable code issues (lint and static-analysis findings — not guesses), proposes concrete fixes, and opens a pull request once a human approves the changes.

  • fix-radar

    Jayesh Agarwal

    Fix Radar closes the loop between code review and code fixing. Qodo reviews a pull request and flags a real issue; a TrueForge agent picks up that finding, reproduces it in an isolated sandbox with its own regression test, writes the smallest correct fix, verifies it, and opens…

  • Rootless

    Onkar Tulshigar

    Domain Hunter finds suspicious domains that may be impersonating a brand. It checks the domains, looks at current and historical evidence, and gives a simple risk assessment for a human to review.

  • ripcord

    Aman

    Ripcord is a simulation-first, on-chain incident response agent for DeFi protocols. In the event of a live exploit (like a reentrancy attack draining a vault), speed is critical but autonomous agents are dangerous if given unchecked write access to production smart contracts.

  • thExplorers

    Saurabh Gupta

    Wrong fixes. Conversion drops, three dashboards show the same number, a PR goes out, and the real cause was a deploy or a log line nobody pulled. LOOP is for the person who owns the product.

  • DevGuard

    Aniket Patel

    DevGuard is a control plane for agentic software development. It lets engineering teams run AI agents that can inspect repositories, modify code, execute changes in an isolated sandbox, open pull requests, and respond to review feedback - all under explicit, enforceable policies…

  • harness-os

    Harsha Priya Ganapathy

    Harness OS is an adversarial crash-test lab for autonomous AI agents. As AI agents move from answering questions to taking real-world actions—issuing refunds, modifying repositories, deploying software, or calling external APIs—they can fail in ways traditional software tests…

  • Sentinel

    Mayank Sharma

    Sentinel is an autonomous AI security investigator that correlates identity and network evidence, produces an auditable threat assessment, and proposes containment.

  • coercion-brake-trueforge

    Vanshika Jadam

    Project Name: Asset-Hostage Shield: The Digital Coercion Brake The Problem: Ransomware attacks and scams weaponize extreme urgency (e.g., a 1-hour countdown to leak assets or pay Bitcoin). This induces panic, driving users or security teams into hasty, unverified decisions.

  • patchforge

    Vivek Silimkhan

    Developers who maintain dependencies. When a vulnerability is found, you need to upgrade, refactor the broken code, test it, and approve it, that's a day's work every time.

  • sentinel-agent

    Prince Panchani

    sentinel-agent is an autonomous incident responder for production software, built on the TrueForge harness. Give it a production incident and it investigates end to end — reads the incident, characterises the symptom, enumerates recent deployments, reads the actual diffs…

Offline

Offline Project Submissions

Handed in live in the room in San Francisco, on August 29, 2026.

55 projects

  • agentlens

    Weigang Geng

    AgentLens is observability for AI agents. Teams running agents on TrueForge have no way to see what their agents are actually doing - which sessions failed, which tools are flaky, where tokens and money go.

  • Aloha Live

    Seth Caldwell

    Aloha Agent connects travelers, locals, and the nonprofits serving the places they visit. Through voluntourism and Aloha Live experiences, our agent continuously researches what’s happening in a community, ingests live signals from travelers, verified locals, nonprofits, and the…

  • itinerary-agent

    John Huang

    Compass is a agentic travel planner. I have traveled a lot, and it always takes hours to research and make the plan. This is my attempt to create an agentic app for consumers.

  • DevGuard

    Aniket Patel

    DevGuard is a control plane for agentic software development. It lets engineering teams run AI agents that can inspect repositories, modify code, execute changes in an isolated sandbox, open pull requests, and respond to review feedback - all under explicit, enforceable policies…

  • hometown-grand-prix

    Karthik Shetty

    Street Steward is a 3D racing game on real OpenStreetMap roads (Downtown San Jose) where AI agents govern the race instead of hard-coded rules. An AI Traffic Steward watches your speed against the real road's speed limit and rules on violations in real time — but the driver must…

  • DreamTeam

    Gracelyn Newhouse

    Agents are overconfident about method. Given an unfamiliar task, they improvise a plausible approach instead of consulting the published research that already solved it, and then report success on checks that never ran.

  • aug29

    Tri Nguyen

    It's for on-call engineers and SREs. The real pain isn't fixing the bug — it's the first ten minutes of an incident spent figuring out what even happened and whether it's your fault or a vendor's.

  • parity-agent

    Bryan Samuel James

    parity-agent detects personalized and dynamic pricing, the kind you can never spot yourself since you only ever see your own price. It fetches the same product page from isolated sessions in different countries and compares what each one is shown, distinguishing real…

  • Agent Fight Club

    MOHAMMED ADNAN

    It's a 3D city where every building is a real GitHub repo. Click one, pick an issue, and two AI agents race to fix it in sandboxed clones while their commits stack up as tower floors, then a referee re-runs the original tests, code quality scores both diffs live, opens the…

  • agent-harness-hackathon

    Akshata Madavi

    Problem: AI evaluation platforms show grader scores and traces, but when a score is low, engineers still have to manually inspect traces, grader logic, prompts, tools, and code to understand why it failed and whether the agent or the evaluator is actually at fault.

  • beforeyoupay

    Lea Frances Yu

    BeforeYouPay is an AI accounts-payable investigator that checks suspicious invoices before a business sends payment. It is designed for accounts-payable teams, finance leaders, operations staff and small businesses that may not have dedicated fraud analysts.

  • Team JC

    Carl Okpala

    We are shipping web software faster than ever, and almost nobody checks whether disabled people can use it. The tools that exist read the code but never use the site. A button can say "closed" in the code, open a menu when you click it, and still say "closed" afterwards.

  • gofer-trace-smb

    Jason Okorie

    Small companies struggle to harness AI agents when it comes to company-specific context. Gofer Trace helps companies give agent harness the context needed to pilot agents across different tasks. Local businesses setup LLM knowlege bases to help give insight on day-to-day tasks

  • PharmaFlow

    Ramya Velaga

    PharmaFlow helps pharmacy and healthcare teams keep track of FDA safety alerts, drug shortages, inventory, and medication availability in one place. Its agents collect live information, work together, and notify teams when something needs attention.

  • Ratchet

    Shryuk Grandhi

    Every step is a git commit plus a sandbox snapshot, and every candidate patch must clear a seven-stage verifier gauntlet (build, cheat check, fail-to-pass, pass-to-pass, types, lint, diff hygiene) before it sticks.

  • Cross Exam

    Martin H. Pefaur

    CROSS-EXAM is a dry run for AI actions that cannot be undone. When an agent proposes an irreversible action — a bulk refund, a payout, closing an account, every guardrail on the market grades the action the agent declared.

  • petassist

    Yu Iwase

    Petassist makes AI agent permission visible to people who are not engineers. Pets live in a pocket window and stick to the monitor like notes. A cat investigates untrusted mail inside TrueForge’s sandbox and cannot send. A bunny drafts a refusal and cannot send.

  • SOLO: Aishow "詠唱" — "incantation" in Japanese.

    Yoshihiro Sekiya

    Aishow ("詠唱" — incantation) is a voice-summoned AI agent for macOS. The problem: as a Japanese engineer at a US startup, every cold outreach through a company's website contact form costs me ~15 minutes — research the company, draft in English, fix the translation, paste it in.

  • ember

    Sviatoslav Zinevich

    I decided on creating an incident resolver "Ember": Submit your critical software error and the agent will identify the error, explain it, and resolve it. From pool exhaustion, worker/thread starvation to poisoned-config, poison-message, stuck queue, and more.

  • Meeting-Action-Agents-

    Vidya Jayaraman

    What does your project do? Meeting Action Agent turns messy meeting transcripts into actionable next steps. It uses a multi-agent workflow to extract decisions, action items, owners, deadlines, and open questions, researches the company discussed using Bright Data, and generates…

  • data-harnesses

    Thomas Barrios

    made just for me to track my macros and health

  • Hush-Harness

    Shradha Pujari

    Hush is an autonomous data-center incident triager and operator. It turns a large alert storm into one evidence-backed root cause, enriches the incident across infrastructure systems, proposes remediation, executes safe actions, and requires a human decision before any…

  • fde-robot-soc-harness

    suhaas

    Physical SOC turns a Reachy Mini desk robot into the human-in-the-loop approval gate for automated vulnerability remediation. Scanners that find CVEs and bots that write patches are solved problems.

  • ForgeRoom

    Pramod Thebe

    ForgeRoom is a shared workspace where people collaborate with persistent AI coworkers—not one-off chatbots, but durable teammates with their own sessions, permissions, and audit trail. Problem: Most AI agent workflows are opaque and hard to trust.

  • countermove

    Sterling Cobb

    Countermove helps you make any sort of business decision by simulating outcomes

  • yallapost

    Taha

    YallaPost is a daily content agent for creators on the startup and venture beat. The hard daily decision is what to make today, so the agent scouts a watchlist of Instagram, X, search, and RSS sources, surfaces three trending topics with clickable evidence, writes a beat-by-beat…

  • ai-video-studio

    AI HOLLYWOOD

    Touchdown AI Video Studio turns a plain-English creative brief into a structured story, six-shot storyboard, production prompts, and an approval-controlled video render plan.

  • mcp-sentinel

    Kiran Devihosur

    mcp-sentinel checks whether an MCP server is safe before your team trusts it. It reads the server's tools, runs the server in a sandbox with fake secret values, calls its tools, and watches whether it tries to steal those secrets or hides instructions in its tool descriptions.

  • Winning Team

    OUSSAMA ELFIGHA

    UpgradePilot turns a dependency-breaking change into a verified pull request. A developer provides a GitHub repository and upgrade target. UpgradePilot retrieves live migration documentation, identifies affected call sites, reproduces the breakage in an isolated sandbox…

  • RoolyTooly

    Raphael Khalid

    - roolytooly is a self-learning harness: an agent that stops repeating its own mistakes. - Coding agents fail in the same families of ways over and over: reporting a stale artifact as fresh, calling a suite green after running one file, claiming "done" on an empty export…

  • Everest G1

    Kaushik Sivakumar

    Everest G1 rescue ops are difficult and takes a lot of time in extreme mountain terrain, where reaching an injured or lost climber can put additional human rescuers at risk.

  • truthlease

    Rahul Krishna

    TruthLease is an agent safety harness that helps operators safely fix previously valid actions when conditions change. It monitors action contexts for stale evidence, re-validates the facts, routes results through a transparent analysis and approval flow, and applies only…

  • whisperer

    chris mckenzie

    It finds out why people are switching to or leaving your product, discovers difficulties they are having, confirm if the bug or issue is real, patches it if it is, follows up with the user and improves the product.

  • Sponde (team name), SOLO

    Abhinav Garg

    Sponde is a deal room for AI agents. Everyone is about to have a personal agent, but there is nowhere safe for two of them to make a deal. In Sponde your agent and my agent negotiate through a neutral room where only validated offers can cross, so private stuff like salary…

  • reweave

    NarasingaMoorthy V

    Every team that extracts data from the web is on the same treadmill: r/webscraping practitioners report 10–15% of scrapers breaking every single week as sites change structure, and one person maxes out maintaining ~100 of them.

  • artiji-buyer-agent

    Robert Schwentker

    The project solves the trust & reliability gap in paid, deferred MCP services. It enables an AI buyer to inspect material terms before payment, obtain human approval, pay exactly once, survive restarts, and verify that the eventual artifact matches the purchase.

  • mcp-breaker

    Dylan Flud

    MCP Breaker is a security testing tool for AI agents that use MCP. Giving an agent access to tools like GitHub, files, databases, or internal services is easy.

  • TEAM BRAHMA

    Sai Nellutla

    BRAHMA is a fault-tolerant command center for autonomous AI workers. It takes a production incident from alert to verified remediation: investigating the issue through MCP tools, dynamically creating specialized AI workers, recovering when workers fail, testing AI-generated…

  • Edit-AI

    Sunny Rodrigues

    EditAI is an open-source video editing agent. You upload a video, tell it what you want in plain English, and it plans the edit, writes its own ffmpeg commands, and runs them in an isolated sandbox.

  • OPESHA Credit Intelligence Agent

    John Wangombe

    OPESHA Credit Intelligence Agent is an agentic AI system designed to help lenders assess SME credit readiness faster and more intelligently. The agent researches an SME, gathers relevant publicly available alternative-data signals, analyzes supplied financial information…

  • MineForge

    Andre de la Cruz

    MineForge lets AI agents operate as embodied workers inside Minecraft, where they can gather resources, hunt, craft, and construct imported blueprints. It gives Minecraft players and agent developers a safe way to experiment with autonomous, observable multi-agent behavior.

  • Incident Responder

    Mayuresh More

    Ripcord is an AI on-call engineer. When a production alert fires it investigates with read-only tools, writes and runs its own diagnostic code in a sandbox to quantify what it found, proposes exactly one remediation — then stops dead and asks a human before doing anything…

  • warren

    Kian Kyars

    Ingests a ChatGPT export, clusters unfinished threads, researches the stale ones against a local corpus, drafts a morning plan, and does not write that plan until a human commits it.

  • scholarly

    Pulkit Arya

    Scholarly

  • hold-the-line

    Karim Baba

    Hold the Line is a phone number you can call. On the other end is a claims agent for a fictional insurer working a vehicle total-loss claim: it pulls the policy, values the car, and computes the settlement. The point is the console beside it.

  • dgx_hack

    Vaibhav Maheshwari

    Agents take a lot of time to do their work and which is why people just run them and let it YOLO stuff, this is great but it creates a problem where the person who gave the assignment for the same does not have the visual feedback for the work done, what we have done is created…

  • Dgx

    ali amjad

    WaterOps is a role-based AI operating layer for industrial water treatment plants. It connects specialized AI agents to an existing Ignition HMI through MCP, with TrueForge acting as the local agent harness.

  • upstream-watch

    Kush Ise

    Upstream Watch reads the things your code depends on — hosted APIs like OpenAI and Stripe, and open-source packages like Express and React — and finds the changes that will break you before your users do, then patches the code, opens a pull request, and stops and waits for a…

  • Nolan

    Aditya Bhat

    Engineering ships faster than marketing can distribute. That's the gap Nolan closes. Product and engineering teams now ship weekly, sometimes daily. Growth and marketing can't produce launch content at that rate: every feature wants a demo video, a LinkedIn post, an X post…

  • agent-hackathon

    Suman Polepaka

    LearnLoop — From content to capability. LearnLoop is a visual, persistent learning agent that turns confusion into a personalized 8–15 minute practice loop: See → Predict → Practice → Prove → Advance Learners can paste a YouTube link, describe where they are stuck, or submit…

  • Team Wingman

    Smaran Aramballi Sandarsh

    Wingman is an AI agent that handles the early stages of dating for you, and stops to ask permission before it says anything you did not authorise. You never fill in a form.

  • placebo

    Hari Gangadharan

    I am a game dev on UGC platforms like Roblox & UEFN, but I face an issue when using AI- it cannot debug games (even with tests, it does not test if it actually made a difference, since the models reward hack!), and I cannot run it locally: - I made Placebo to solve this, since…

  • trueforge-agent-harness-hackathon

    David Van Gheel

    It is for people looking for jobs. The AI will prompt the user for information about what their skills are, what they are looking for an other relevant information. Then it will give you a list of jobs to apply for and help you to adapt your resume to fit the jobs.

  • BizForge

    Yaroslav Volovich

    BizForge is agentic system which helps entrepreneurs find business or product opportunities. 1) First "setup" agent interviews the user (could be just LinkedIn link) collecting user competencies and business preferences (eg avoid military).

  • two-fold

    Shubham Shinde

    Twofold is a website audit for the two audiences every public site already has: people and AI agents. Today, teams measure humans with analytics, heatmaps, and A/B tests. Agent visitors — copilots, research agents, scrapers — hit the same pages, fail silently, and leave no score.