
The Agent Harness Hackathon
Winners
Grand Winners
The best of everything handed in, online and in the room.
Best Use of TrueForge
NVIDIA DGX Spark
ONCALL
Elijah Umana
An on-call agent that takes an incident as far as a rollback, with every step of that rollback written down before it happens. The revert it expects is checkpointed before the push, a run interrupted halfway reconciles against the commit the remote actually has, and an incident survives a restart of the harness underneath it.
Best Code Quality
Mac Mini
CapyGuard
Nelson Lai
A guard that runs untrusted input in a sandbox and reports only what it watched happen there. Reading an artifact and ruling on it are split across separate agents, so the text under judgement cannot instruct the agent writing the judgement: a boundary the team found they had enforced by prompt alone, and rebuilt structurally, when their own review trail caught it.
Best Use of Bright Data
AirPods 4
AgentRadar
Team AgentRadar · Biswajeet Sahoo, Priyanshu Solanki
An agent that treats a code review finding as a hypothesis rather than a verdict. It maps the finding onto the code graph, picks the tests that reach the line it names, runs them, and files the result as confirmed, unreproduced, uncovered, inconclusive or unlocatable.
Best blog post
Keychron Keyboard
Mayday
Anamika Singh
An incident responder that does the first twenty minutes of the work and none of the deciding. It pulls the metrics, reads the logs, checks what shipped and arrives with a diagnosis, the evidence behind it and the reasoning shown, then waits for a human before anything changes state. Recovery is judged on the latest raw telemetry rather than on an average the outage is still dragging down.
Online
Online Project Submissions
Handed in from anywhere between August 24 and 30, 2026.
286 projects
failsafe-trueforge
Charannyan Kannan
FailSafe gives an incident responder one narrow job: diagnose a checkout outage, ask for approval for the safest repair, and prove the system recovered. In the demo, an rc3 database-pool change pushes six checkout replicas toward a 200-connection Postgres ceiling.
Tech Titan
Riju Paul
Our project is an AI-powered Hand Matrix system that uses AI to analyze hand-related input and provide useful information or recommendations. It is designed to make the analysis process simple, fast, and user-friendly.
SchemaSentinel
Mohit Pargaie
SchemaSentinel is an agentic database migration safety system for developers and database engineers. It takes a proposed PostgreSQL migration, inspects the actual target schema, analyzes risks such as table rewrites and locking, validates the migration in an isolated PGlite…
Smart-E-Commerce-Support-Refund-Agent
Shaik Mullasadik
Smart E-Commerce Support & Refund Agent is an autonomous AI agent designed for online retail businesses to handle customer inquiries about missing packages and automate the refund process securely.In traditional e-commerce, customer support tokens for "order not received" take…
Scrap-verse-hackathon-
Saniya Khan
Focused on you data and codes
repomedic
ANIRUDDHA ADAK
RepoMedic is an autonomous open-source repository triage agent built on TrueForge. It scans repositories for real problems such as failing CI, broken README links, stale issues, and vulnerable dependencies.
SmartCart_AI-RAG
DURGA PRASAD TATAPUD
SmartCart is an AI-powered multilingual shopping assistant that helps users find and compare products using natural-language queries. It understands English, Telugu script, and Roman Telugu, extracts requirements such as budget, RAM, storage, processor, display, brand, and use…
trueforge-agent-assistant
Dipayan Mukherjee
SafeGuard Agent is an autonomous workspace assistant designed for developers and system administrators who need to securely audit directories, manage local files, and execute maintenance routines without risking security compromises.
AI-Incident-Responder
Sachin Kumar
The AI Incident Responder is an autonomous agent that handles production incidents end-to-end. When an alert fires (e.g., error rate spike on a service), the agent automatically investigates by querying service health metrics and deployment history, spawns parallel sub-agents to…
debrief-agent
Yashaswini.V
Debrief is a mission-dossier style web application that acts as a strict fail-safe for autonomous AI coding agents. Currently, allowing AI agents to navigate repositories is dangerous—they operate as black boxes that can blindly push broken code, leaving developers to clean up…
Poorva Shrivas
Poorva Shrivas
It is a safe agent on TrueForge that uses tools safely and can be stopped - solves unsafe agents problem for teams who need job-ready agents
KARAN YADAV
Karan
Projection Witness is a proof-carrying repair agent for teams operating event-sourced applications on PostgreSQL. It solves a subtle production failure where an out-of-order event is permanently skipped by a projector.
sentinelforge
Aayush Sahu
SentinelForge is a safety- AI incident responder designed for engineering teams. SentinelForge looks at repository incidents by using read- evidence. SentinelForge finds the root cause. SentinelForge creates a repair proposal that is limited in scope and gives it a fingerprint.
Aegis
Bradley
Aegis is an incident response platform for blockchain protocols. When something goes wrong on chain, like a treasury draining below the gas it needs to keep paying out, a contract acting strangely, or a big unexpected transfer, Aegis investigates on its own and comes back with a…
licence-to-patch
Mayank Pande
What it does: Licence to Patch is an agent that reviews your dependency-update pull requests and tells you which bumps you can trust — then, only with your approval, fixes the one you can't.
kernel-preflight
Swapnil Dutta
kernel-preflight is a GPU kernel-writing agent that cannot lie about the speedup — it submits kernel source and never submits a number. An agent that writes a kernel and reports its own speedup has both the incentive and the opportunity to be wrong in its own favour.
jarvis-ops-agent
Sahil Mane
Jarvis is a voice-first personal operations agent. You give it an outcome, not a sequence of app commands. It decides what evidence it needs across connected systems, does the work, and stops for human approval before anything irreversible.
divvy-forge
Hitesh Pattanayak
divvy-forge automates the research-and-draft step for dividend portfolio reviews. It reads my existing dividend holdings from a GitHub portfolio tracker, fetches live fundamentals (yield, payout ratio, FCF, EPS) from Screener.in and yfinance, runs two parallel subagents — one…
k8s-sentinel
Anshul Bisht
Autonomous Kubernetes incident triage agent built on TrueForge — real MCP tools, sandboxed analysis, subagent parallel dives, approval-gated remediation.
contain
Deep Shah
ContAIn is a secret-leak remediation agent. When a credential gets committed to a repository, someone has to work out three things by hand: does it still work, what can it reach, and will killing it break production.
casework
Nick Sawinyh
Casework works a transit data steward's feed-failure queue. A monitoring dashboard for a state's GTFS feeds shows dozens of failing feeds, but the person on rotation has a different question: how many problems is this really, and who do I write to?
arcadeops-mission-control
Damien CREDOZ
ArcadeOps Mission Control is a proof-carrying incident-response agent for operators who need AI to act without losing human control. It investigates a fictional degraded checkout service, delegates independent verification, validates a rollback in a Daytona sandbox, pauses for…
TrueForge
Chirag Tankan
AutoCompliance-Agent is an autonomous DevSecOps and regulatory compliance engineering agent designed to bridge the gap between compliance audits and code-level remediation.
FaultTrace
Harshul Dwivedi
FaultTrace is an autonomous forensic agent for vehicle diagnostics. Given one trouble code (e.g. P0171 "System Too Lean"), it runs a full investigation: gather evidence, enumerate competing hypotheses, compute a deterministic Bayesian differential, and spawn per-hypothesis…
checkout-services
Sourjya Saha
When production microservices crash in the middle of the night, traditional AI chatbots can only offer theoretical advice to tired on-call engineers. They cannot safely investigate, reproduce, or patch the live codebase without risking production outages.
Patchwork-AI
Arpita Sharma
Patchwork-AI is an autonomous developer agent designed to automatically detect code issues, simulate non-mutating patches, and streamline code reviews to maintain high software development standards.
auto-patch-agent
Sai Rathod
Auto-Patch Agent is an AI-powered incident response agent for production code failures. It takes an incident such as an HTTP 500 error and investigates the relevant repository code to identify the root cause.
DHARM PARESH KOSHIYA
DHARM KOSHIYA
SentryOps is an autonomous incident response agent built to monitor system alerts, inspect error traces via MCP tools, run automated regression tests in an isolated sandbox, and draft rollback patches.
Double-O-Harness
Abhay Singh
Double-O-Harness is an ultra-cost efficient, autonomous Open Source Intelligence (OSINT) agent capable of executing hundreds of concurrent web searches and compiling multi-page PDF dossiers with images.
Licence
Leo
Licence is a Harbor Pay on-call desk on TrueForge. Checkout error rate jumped from 0.4% to 8.1% after deploy 4c21 (PAGER-4419). The agent reaches a real MCP connector, fans out subagents across metrics, deploys, and logs, bisects the last four deploys in an isolate, and stops.
undox
Manas Dutta
Undox finds people-search sites that are leaking someone’s PII and helps drive opt-outs — but it never submits until a human allows the exact payload. It’s for anyone who cares about privacy, because removal today is slow, easy to mess up, and hard to undo once you hit Submit.
catherine (rin) pereira
catherine (rin) pereira
lifeagent is an approval-gated personal routine manager that reads your google calendar, sends push notifications for classes and events, reminds you to take medications at 8am and 8pm, nudges you toward sleep at 11pm, and runs a laundry countdown timer.
dossier
Sanjeev Kumar S
Dossier is an adversarial multi-agent idea intelligence platform that stress-tests startup concepts and technical proposals to find fatal flaws before capital or engineering effort is spent. Solves LLM sycophancy and ungrounded optimism.
evidenceforge
cmdr-chara
EvidenceForge is an evidence-gated control plane for investigating and repairing failed CI workflows. It binds every incident to an exact repository revision, coordinates three read-only diagnostic specialists, reproduces the failure in a Daytona sandbox, verifies the patch…
treasuryforge
Romanch Roshan Singh
TreasuryForge is an autonomous treasury agent that manages a simulated portfolio across cash, crypto (BTC/ETH), and NSE equities, built natively on TrueForge.
Ship_Safe
Dhruvil Vora
ShipSafe is an agentic PR test runner and code review assistant designed to make pull request validation safer and more automated. When given a pull request, ShipSafe analyzes the changes, coordinates agents to review the code, and runs relevant tests inside an isolated…
Sparkles
Harieesh Ragaw SK
ReleaseGuard is an AI-powered Release Safety Engineer built on the TrueForge agent harness. It autonomously analyzes software changes by inspecting pull requests and CI results through GitHub MCP, understanding repository architecture with DeepWiki, investigating production…
RoboOps
Al Mahmud Samiul
RoboOps is an autonomous AI Robotics Site Reliability Engineer (SRE) that diagnoses and resolves Autonomous Mobile Robot (AMR) fleet failures using TrueForge and MCP.
msdesk
Anvit Devadiga
Operations teams, engineering managers, and support leads waste hours every week manually pulling raw support metrics, calculating KPIs, and formatting status updates for Slack channels.
PriceMCP
Tobias
PriceMCP gives AI agents one MCP interface for trustworthy price comparison. It resolves an exact product variant before ranking normalized offers, then returns merchant identity, price, currency, freshness, availability, conditions, and provenance.
SentriX
Anish Das
SentriX AI solves the problem of finding and fixing security vulnerabilities in modern codebases. Traditional security tools often identify issues but leave developers to manually analyze and fix them.
Battery_Life_Prediction-Cost_Optimisation
Anik Bera
Modern battery management and manufacturing face three critical bottlenecks: The "Knee-Point" Degradation Trap: Batteries degrade non-linearly. Traditional battery management systems (BMS) rely on Coulomb counting or simple linear decay, causing them to miss the sudden…
leadforge
Dharanidhara D J
LeadForge finds leads and calls them to close deals. Find. Call. Sell without the huge manual work.
sentinelforge
Navraj Singh
SentinelForge is an autonomous Site Reliability Engineering (SRE) incident responder for on-call engineers and DevOps teams. When a production service breaks at 2:00 AM (like a checkout API returning 504 Gateway Timeouts), engineers usually have to wake up, manually comb through…
ForgeGuard
Pranav Gupta
ForgeGuard is a tool that checks whether an AI agent does tasks safely and correctly, not just whether it gets the right final answer. For example, if an agent is told to fix failing tests, it might actually fix the bug, while another agent could simply delete the tests.
Ctrl C + Ctrl V
Harshit Saxena
RefundGuard is an AI-powered refund investigation agent for customer support teams. It uses TrueForge and MCP tools to investigate customers, orders, payments, refund history, and policies, then recommends a refund decision.
TrueSRE
Vishal Rajput
TrueSRE is an enterprise-grade autonomous Site Reliability Engineering (SRE) multi-agent platform that detects, diagnoses, dry-runs, and remediates production Kubernetes outages in under 25 seconds—reducing Mean Time To Resolution (MTTR) from 45+ minutes.
self-correcting-integration-maintainer
Keniel Maldonado
Most code review catches what a rule already describes. The failures that actually ship are the ones nobody wrote a rule for. This is a Self-Correcting Integration Maintainer: an agent that inspects code it has never seen, is told nothing about what is wrong — the word "bug"…
GOOSEGOOSE
Sami El-Figha
The Gaggle is an adversarial multi-agent system for microbiome R&D. Instead of trusting a single AI model to produce a confident answer, it brings together independent AI scientists that argue for and against candidate hypotheses, retrieve live scientific evidence, run…
peel
Srinivas Sivaratri
Peel is a safety layer for sending .xlsx attachments from an owned mailbox. It preserves the exact attachment and SHA-256 hash, checks the workbook for hidden or unsupported content, runs verification in a fresh Daytona sandbox, and requires separate human approval before the…
Raghav Negi
Raghav Negi
ForgeOps is a terminal-native CLI agent for code review and incident debugging. It connects to a pre-configured TrueForge agent (forgeopsv1s) from any terminal - VS Code, GitHub Codespaces, Windows CMD, or Linux — and lets developers interact with their GitHub repositories using…
PMKVY
SAI DUTTA ABHISHEK DASH
SENTRY is an on-call incident responder that runs on TrueForge. When a production alert fires -- payment failures, 5xx spike, latency climb -- it investigates end-to-end over MCP: reads Prometheus metrics via Grafana, checks GitHub deploy history, correlates cause, and proposes…
forgeos-lite-trueforge
Edson Dasilva
ForgeOS Lite is a safety-first AI coding control plane for developers who want autonomous agents to build real software without giving them unrestricted authority over the original project. A user describes an idea in natural language.
sre-oncall
Nitish Mane
Two agents on the TrueForge harness. sre-oncall receives a Grafana alert, investigates over live MCP tools, correlates the failure to the ArgoCD sync and commit that caused it, and opens a revert PR carrying its evidence; a human merges, ArgoCD syncs, the alert clears, and it…
clinical-scrubber
Biplab Bera
Clinical Trial Data Scrubber & Analyst. Medical researchers spend hundreds of hours de-identifying patient data and running statistics before they can draft an FDA report.
mhcvi-live
Shourya Sharan
The problem no one talks about When disaster funding gets allocated, the algorithm picks winners by cost-per-person. Urban clinics get funded. Remote island communities get skipped. Not because they don't need help — because they're expensive to reach.
AI-Employ-Management-System
Kajal
Employee Management System is a web-based application that helps organizations manage employee information efficiently. It allows users to add, update, view, and delete employee records, making employee data management faster, easier, and more organized.
fixforge
Ananya Unmesh Chaudhari
FixForge is an autonomous software debugging agent. Given a bug report and a codebase, it investigates the actual project — listing and reading relevant files, reasoning about the root cause — then proposes a code fix.
AFTERMIND-WEBSITE
MOHAMMED SAFIA TABASSUM
What does your project do? AfterMind is an AI-powered emotional storytelling platform that transforms a user's personal thoughts, experiences, and emotions into meaningful visual stories.
clarity-agent-hackathon
Ghulam Mustafa
Clarity is an approval-first AI agent workspace that gives AI models a license to act safely under complete human control. It solves the risk of unchecked autonomous agents by introducing a Human-in-the-Loop safety layer.
airlock-mcp
Himanshu Kumar
Airlock audits what an MCP server actually does before an agent trusts what the server says about itself. It inventories declared tools, exercises each under a capped probe budget and compares observed behavior with published annotations.
verdict
Himanshu Kumar
Verdict turns intermittent GitHub bug reports into checkable reproduction records. Three bounded agents investigate the trigger, localize only what the evidence supports and design a regression plan.
Team Roferzzz
Aishwary Gupta
TrustForge is an autonomous AI Web3 auditor. It uses a TrueForge multi-agent system and a custom Python MCP server to automatically analyze Solidity smart contracts for vulnerabilities.
agent-change-gate
Huang Chung Yi
Agent Change Gate tests an agent instruction change before it reaches GitHub. It runs the current and candidate specs against the same 16 content-hashed scenarios five times, separates unstable results from regressions, then asks a person whether to create the branch, commit…
aegis-zero
Aryan
Aegis Zero is an autonomous DevSecOps agent that scans Python code for security vulnerabilities (SQL injection, hardcoded secrets, command injection), generates secure patches, and commits fixes to GitHub — but only after human approval.
true-trade
Mayank Senani
true-trade is a paper-trading agent for NSE stocks that runs on TrueForge. It screens a universe of 100 Indian large caps using free yfinance data, picks momentum candidates, and trades virtual money on a simulated clock — you start it on any past date, it trades on only the…
PaywallProof
Vasu Bansal
PaywallProof checks that a subscription provider's state matches the access a SaaS app actually grants. It runs four scenarios across API, browser, and application state: free access, paid access, scheduled cancellation, and post-expiry denial.
sentinental
Mayank kumar upadhyay
SENTINEL is an autonomous supply-chain security analyst built on the TrueForge agent harness. Point it at a repository and it takes dependency-vulnerability triage off your plate entirely — right up to the point where a human should decide. The problem. A CVE advisory lands.
secureops-guardian
Jayesh Savaliya
SecureOps Guardian is a human-controlled security incident-response agent for on-call, DevSecOps, and platform engineers. It connects a security alert to the exact GitHub change that caused it, checks the evidence, and prepares a safe, least-privilege fix.
SafeOps
Sahil Warade
SafeOps is an autonomous AI incident-response agent built on the core philosophy: "Investigate autonomously. Act only with permission." During critical production outages, on-call developers and SREs lose valuable time manually correlating metrics, deployment logs, and code…
dilate
Devansh
Dilate is an autonomous QA engineer for time-dependent bugs. It scans a codebase for logic involving dates, time zones, expirations, schedules, daylight saving time, and calendar arithmetic, then automatically generates targeted test scenarios and executes the application under…
APIx Predators
Gautam Khosla
Keyring is an access governance agent. You give it one instruction, such as "audit access for Ada Lovelace" or "offboard this contractor", and it fans out across your connected systems to inventory every grant a person or a machine still holds.
Parth Pinjarkar
Parth Shrikant Pinjarkar
AutoVault is an autonomous AI security operations agent that detects, investigates, and responds to ransomware attacks in real time — powered by TrueForge. The Problem It Solves Ransomware costs $20B+ yearly.
Orion Forge
Arul Raj W
Our UAV Disaster Management System is an AI-powered platform designed to assist emergency response teams during disasters. It analyzes UAV/drone footage using computer vision to detect people in affected areas and provides location-based information through an interactive…
sma-calibration-agent
Juniad
NEUTRINO identifies the seven parameters of a superelastic shape-memory-alloy (NiTi) constitutive model from a noisy tensile test. This is an inverse problem.
AgentHarness-Truefoundry-hackathon
Damrukesh Daliparti
An agent for a fictional CI/CD tool that answers support questions from a knowledge base, files GitHub issues, and reviews pull requests against documented conventions. Every write action (creating an issue, posting a review comment) pauses for human approval first.
codepulse-ai-agent
Puneet bhardwaj
CodePulse AI is an autonomous agent that fixes technical debt and broken test suites. It automatically scans repositories, isolates environments in a Daytona sandbox, applies patches using NVIDIA Nemotron, and opens verified GitHub Pull Requests for developers.
trueforge-incident-agent
Bhagyashri Sinnarkar
TrueForge Incident Response Agent helps engineers quickly investigate and resolve system incidents. It automatically collects logs, metrics, and incident data, analyzes the problem, and suggests possible solutions.
spyglass
Vivek
Spyglass is an evidence-engineering layer for AI incident-investigation agents: a Rust engine mines production logs, metrics, and deploy events into ranked, bounded, cited evidence bundles (served over MCP) instead of letting the agent page through raw telemetry then verifies…
trueforge-tally
Vishal Jha
It basically provides Indian SME's and MSME's to connect their accounting software like Tally directly to their agent. The navigation of tally is such a laborious task that it takes an CA 10 month just to master the software.
Circuit_Breaker_Agent_H
Sandeep
Circuit Breaker is a deterministic authorization control plane for AI-powered financial agents. It sits between an AI agent and the payment execution layer.
headcount
Sarthak Agrawal
HEADCOUNT is an idle game whose economy is human attention — and an AI agent designs the game while you approve its changes. Idle games model a company where labour is perfect: a machine makes 3.2 units a second forever, unsupervised.
Git Going
Anugrah K S
ARCHAEOLOGIST — Digital Forensic Workstation is an AI-powered agentic forensics and archaeology tool designed to reconstruct the context, architecture, and intent of abandoned or undocumented software projects.
PatchForge
Praveen Rajan
PatchForge is an agentic security platform that automates the end-to-end process of identifying, testing, and fixing software vulnerabilities (CVEs) within isolated, sandboxed environments.
trueforge-security-agent
Darshan S
TrueForge Security Agent is an autonomous security tool for enterprise data sanitization and forensic file recovery. It scans storage drives and soft-deleted recycle bin manifests, uses AI content inspection to classify sensitive data (SSNs, credit card numbers, API credentials…
Nutri-Ninja-5.0
Jitendra Dewangan
An AI-powered food scanner Android app. Scan any packaged food barcode, get instant health scores, ingredient warnings, personalized nutrition advice, and AI-powered diet chat — all stored locally on your phone.
Curious Pixels
Tejas Krushna Rawool
OpsForge is an AI-powered SRE command center that helps teams respond to production incidents faster. It automatically investigates logs, metrics, and deployments, suggests safe fixes, and verifies the results all while keeping humans in control of risky actions.
deadman
Daniyal Ahmed
DEADMAN is an AI SRE that doesn't just diagnose incidents it fixes production and undoes its own mistakes. It's for on-call engineers drowning in 3am alerts: it investigates from live signals, applies the fix, verifies it held, and auto-rolls-back if it didn't.
Crucible
Bhagyesh Kamal
The Crucible is a safety-first autonomous security validation agent. Given a controlled vulnerable target, it investigates, forms a vulnerability hypothesis, writes and tests its own PoC, and only after a human authorises the one sensitive step runs a controlled exploit to…
AEGIS-007-Autonomous-Incident-Commander-and-Harness
Raj Tiwari
Problem & Target Audience: A chatbot answers questions, but an autonomous SRE agent acts on production infrastructure: querying metrics, bisecting repositories, and executing rollbacks.
repo-debugger
Shaik Nakeeb Mufeed
Repo Debugger is an AI agent that helps developers fix bugs in GitHub repositories. Developers paste a repo URL and error description — the agent inspects the codebase using real GitHub tools, identifies the root cause with evidence from the actual code, proposes a specific fix…
bumpsmith
Aryan Gorde
bumpsmith migrates a Python codebase from pydantic v1 to v2, and it only keeps an edit that a green test run stands behind. The problem is narrower than "migration is hard".
Wemakedevs
Golla Satvik Kumar
Data Watchdog is an agent that checks a database for underperforming categories, judges whether the finding actually crosses a significance threshold, and only escalates to filing a GitHub issue when it does — every sensitive action, from running the SQL query to creating the…
RE-Agent
Shubham Yadav
RE Agent analyzes unknown Linux binaries safely — it runs static and dynamic analysis inside a hardened Docker sandbox, drafts a triage report, and pauses to ask for human approval before publishing anything to GitHub as a pull request.
auditforge
Priyanshu
AuditForge is an autonomous cybersecurity and DevOps incident triage agent that provides cryptographic audit trails for every decision it makes. When high-severity incidents strike (such as connection pool exhaustion, unmasked credential exposure, or CI/CD regressions)…
reroute-lg
Ansuman Satapathy
ReRoute-LG is an autonomous agent that triages supply-chain disruptions. When a typhoon or port strike threatens a shipping corridor, it checks live weather and news data, calculates how many days of inventory are left, evaluates alternate suppliers against cost, lead time, and…
trueforge-secure-agent-harness
CHANDRA PAVANSAI
The project, "The Safe AI File Guard: TrueForge Interception Harness," is an enterprise-grade incident response virtualization layer designed to allow autonomous AI agents a safe runtime environment.
handler
Udit Jain
HANDLER is the on-call engineer for someone else's training runs. A run dies at 03:14 and nobody notices until 09:00 — the GPU bills for six hours of producing nothing, and whoever owns the run starts their day reconstructing what happened from a log file.
Right2Erase
Senali
What does project do? Right2Erase is an AI-assisted system for handling GDPR “right to erasure” requests. It investigates a customer’s data across databases, file storage, billing systems, and logs, builds a deletion plan, and rehearses the plan in a disposable sandbox.
Sentinel
Ronak Paul
Sentinel is an AI SRE incident-response platform. When a production alert fires (via webhook from Grafana/PagerDuty) (We can create webhooks per project from the sentinel portal), it investigates root cause, verifies fixes in an isolated Daytona sandbox, and only ships changes…
deployguard-ai
Anshu Gupta
DeployGuard is an AI incident response agent for production and DevOps teams. It investigates service incidents by checking metrics, logs, and recent deployments, then identifies the likely root cause and recommends a fix.
SD007 - SOLO
Sri Dharshan G D
Odezzy AI audits and governs the MCP (Model Context Protocol) tools an AI agent can call. It scans every connected MCP server's tool definitions for prompt-injection patterns, leaked secrets, and schema mismatches; runs an LLM-based semantic check and embedding-based drift…
digital-declutter-agent
Evangelin Blessy H K
Digital Declutter Agent is an AI-powered filesystem assistant that investigates messy folders using MCP tools. It can recursively scan directories, identify duplicate files, read supported text and PDF files, process images through MCP ImageContent, summarize relevant documents…
codepulse-agent-ui
Puneet
CodePulse AI automates repository bug fixes in isolated Daytona sandboxes. It streams real-time terminal logs and enforces human-in-the-loop pull request approvals to safely accelerate software developer workflows.
hackathon
Savar Bhasin
I've built, The Squad, which is a control center for running a fleet of specialist agents. You give the squad lead a mission, it breaks those down into dependent tasks for agents such as software engineer, researchers, and work intelligence managers.
Proof of Chaos
Mohit Kushwaha
BRANCHPOINT is a safety and control layer for autonomous AI agents that need to take consequential real-world actions. Instead of letting an agent predict what will happen and immediately act, BRANCHPOINT makes it rehearse the action first.
portfolio-agent
Zoya Muskaan
It's an portfolio agent that turns your resume and GitHub username into an actual working portfolio site. Instead of just trusting whatever your resume says, it goes and checks your real GitHub repos first, writes the copy based on what it actually finds, builds a real site, and…
SentinelForge
Akash
SentinelForge is an AI SOC Incident Commander that autonomously investigates security alerts. It delegates work to Log, Network, and Malware investigator agents, gathers evidence through MCP tools, analyzes suspicious PowerShell artifacts safely, correlates findings, calculates…
bomb-squad-agent
Anurag Thopate
When production servers run out of memory or get stuck in cache deadlocks, SREs and backend teams need to fix it quickly. But giving an AI bot direct root access is dangerous because one wrong wildcard or bad command can delete active user sessions or important system files.
migration-sentinel
Sumit S Chawla
Migration Sentinel is an AI agent that investigates database migrations, dry-runs them in an isolated shadow sandbox, validates rollback, and gates production deployment behind human approval — evidence over predictions.
Fifty Shades of Agent
Ashaya Sah
For a Forex/Crypto trader, the project helps gather the technical data, technical analysis, news data and sentimental analysis all bound by Trueforge configured agent with agent harness and executing the trades if confirmed or instructed by a trader.
retrohost
Sameer Goyal
Published results are rarely re-run, so broken code goes unnoticed. Retrohost re-runs a paper's figures in isolated sandboxes via parallel subagents, classifying each REPRODUCED, PARTIAL or FAILED, then asking a human before publishing.
gatekeeper-trueforge-hackathon-
vaishnavi adepu
Gatekeeper is an approval-gated dependency upgrade agent for developers and small teams who know their dependencies are out of date but keep putting it off, because upgrading is tedious and you cannot tell what is actually safe without running something.
CampaignGuard
Savanth JR
My project "Campaign Guard" prevents broken ad funnels and unauthorized production overwrites by AI agents. When a marketing team wants to launch a new campaign, the agent autonomously verifies the destination URL health (preventing 404s) and generates sanitized short-links.
Jstreet
Hemang Varshney
MCP Vetter Agent is the security auditor of third-party MCP servers, running on TrueForge. It validates repos individually prior to agents connecting to them. The agent performs parallel probes (static AST analysis and dynamic).
Outsmart
Harsh Singh
Outsmart upgrades your npm dependencies and fixes what the upgrade breaks. Detection is a solved problem — Dependabot, npm audit and OSV all find the advisory. Repair isn't.
trueforge-hackathon
Abhiraj kumar
Anahita is a Hinglish-speaking, gothic-persona AI junior developer built on TrueForge. You give her a coding task by voice or text, and she works it the way a real junior dev would: she plans the change, writes and runs tests in TrueForge's sandbox, opens a PR, and waits for…
interviewos
Priyam Srivastava
InterviewOS is a research desk for one job: investigate a target software-engineering role, write a cited evidence ledger, compare it to a resume, and publish a four-file prep workspace (evidence.jsonl, brief.md, gaps.md, plan.md) only after a human approves.
Astella
Naitik Morajkar
Consensus Gap is an agent that takes any topic and researches two things in parallel: what experts actually say (from academic, institutional, and scientific sources) and what the public actually believes (from forums, social sentiment, and survey data).
AGENT
Muhammed Nihal P n
The Incident-Responder MCP Agent is designed to solve the critical problem of production outages that cost companies thousands of dollars every minute, serving as an essential tool for engineering and Site Reliability Engineering (SRE) teams.
passage
Tushar Kumar Shahi
Passage automates the health insurance pre-authorization process for hospital insurance coordinators in India. Right now, before a patient can get cashless treatment at a network hospital, a coordinator has to manually look up ICD-10 diagnosis codes, find the right procedure…
repo-rescue-agent
Arthur Yousif
Repo Rescue Agent helps developers diagnose and repair small repository issues. It inspects a repository, identifies a concrete bug, proposes a focused fix, runs verification, and produces a reviewable change for the developer.
migration-mate
Sreeram Reddy Velagala
MigrationMate is a safe, human-in-the-loop database migration agent built on TrueForge. Users describe database changes in plain English; the agent plans, tests in an isolated sandbox, asks for approval, opens a Qodo-reviewed GitHub PR with rollback scripts, merges, executes…
Edith-Hackathon-TrueForge
Aarya Singh
Our project is a multimodal AI desktop assistant designed to help users perform research, coding, productivity, and computer-based tasks from a single interface.
db-rag-agent
Aman Kushwaha
Deploy Detective (D2) is an approval-gated agent that mirrors an on-call engineer across two workflows: Incident response: Pulls real data via MCP, investigates deploys/alerts/metrics in parallel, bisects the culprit deploy in a sandbox, then pauses for human approval before…
AgentX
Dinesh Jinjala
AI agents with database access have a problem: the only thing stopping them from exporting customer emails is a system prompt asking them not to. Prompt injection, long contexts, or one bad completion and the data is out.
ARMAMENT
Moazzam
ARMAMENT is an on-call incident responder. When a service breaks, it does the investigation a human would do at 3 AM — reads container health, searches logs, checks memory and CPU, forms a diagnosis — and then stops.
breakglass
Abhimanyu Goyal
BREAKGLASS is an autonomous red-team security assessment agent that analyzes a target repository, discovers security-relevant indicators and attack surfaces, and generates actionable vulnerability hypotheses.
autofix-agent
Priyanshu Singh
AutoFix-Agent is an autonomous CI/CD failure auto-remediation and governance engine built for software engineering teams. Modern engineering teams lose 20% of their sprint velocity manually triaging broken CI/CD pipelines, deciphering runner stack traces, and writing regression…
student-performance-predictor
Battena Anjali
StudentAI is an AI-powered student performance prediction system that analyzes academic factors such as attendance, study hours, previous marks, assignment scores, internal marks, GPA, and backlogs to predict a student's expected performance.
Harmness
Yash Gupta
Portfolio Desk measures an investment portfolio against its target allocation, works out how far each holding has drifted, and proposes the trades that would bring it back into line.
4Sum
Chirag W
DriftFix is an AI agent that turns breaking SDK changes into tested migration pull requests. In our demo, it upgrades a Python project from Stripe 14.3.0 to Stripe 15.6.0.
sentinelops-control-tower
Ankit Anand
SentinelOps Control Tower is an AI-native incident-response copilot. Given a security/operations alert (e.g. "payment failures spiking in checkout"), it: Investigates autonomously through read-only observability, deployment, and runbook tools (error rates, latency, logs, recent…
stores-triage
Anmol Gupta
Stores Triage is for a stores officer at a locomotive works. A spare part drops below its reorder level and an alert fires; he then has to answer a question the alert cannot — is this real?
Raphael-autonomous-engineer
Satvik Mishra
Raphael is an autonomous junior software engineer built on TrueForge, that reads a GitHub issue, implements the fix, validates it against tests, and waits for human approval before taking any irreversible action.
ColdCall
Dhruv Mulay
ColdCall is an incident commander for pharmaceutical cold-chain temperature excursions. When a shipment of medicine strays outside its labeled temperature range, someone must decide: release it, quarantine and retest it, or destroy it — an industry losing ~$35B a year to failed…
sentinel
Jyotishman Pathak
Sentinel is an autonomous Site Reliability Engineering (SRE) control plane and multi-agent incident management platform for distributed microservices. During production outages, on-call engineers are often overwhelmed by cascading telemetry noise and delayed root-cause analysis.
forty-two
Pawandeep Singh
Forty Two is an AI data analyst that lets users connect databases or upload CSV/Excel files, then ask questions in natural language. It autonomously discovers schemas, runs safe queries, analyzes data in a sandbox, and produces verified answers with durable tables and charts.
Outreach
Pabitra Maity
Outreach is an AI agent for cold outreach. You give it a company, role, and recipient - it researches the company live, drafts a personalized email, and creates a real Gmail draft once you approve it.
Dossier
Bhargava Gumpula
It applies to jobs for you and stops before sending anything. Its for anyone who is trying to apply for jobs. You name a company, and it reads their real job board and ranks every role against your resume, so a board with hundreds of openings becomes the handful you actually…
attest
Kartik Shreekumar
Attest checks whether an MCP server actually behaves the way it says it does. Every MCP tool can declare hints like readOnlyHint: true. Agents rely on those hints to decide what's safe to call without asking a human first. The problem is that nothing verifies them.
the-verifier
Reet Singh
The Verifier checks public claims by researching both supporting and conflicting evidence. It compares source dates, explains its conclusion, and requires approval before saving or exporting a dossier.
cloud-incident-responder
Mrittiga M
Cloud Incident Responder is an autonomous DevOps agent designed to help identify and respond to cloud/server incidents. It reads incident logs, analyzes the reported failure, identifies a likely root cause, and proposes a remediation action.
crucible
shankar
CRUCIBLE is an autonomous Site Reliability Engineering (SRE) and Incident Response platform that detects, diagnoses, sandbox-verifies, and remediates distributed microservice failure cascades.
BBBP-B3DB-Agent
Biplab Ghosh
NeuroPerm AI is an autonomous computational pharmacology and cheminformatics agent designed to screen Blood-Brain Barrier (BBB) permeability for Central Nervous System (CNS) drug discovery.
STRAW HATS
Parth Gupta
Sentinel is a prompt regression gate. When a PR changes a judging prompt, the TrueForge agent reads both versions through the GitHub MCP, runs a 10-case eval suite (subagents in the sandbox, scored with the sentinel-scoring skill), comments the comparison, and stops.
ByteBridge
Simranjit Kaur
CyberForge is an AI-powered SOC agent that investigates security incidents, correlates evidence, recommends containment actions, and requires human approval before taking action.
thinkfill
Gopinath V
ThinkFill is an AI-powered form filling assistant built for the TrueForge Agent Harness Hackathon. Filling forms repeatedly is a frustrating task, especially when users have to enter the same personal information again and again.
DataForge
Mohith Tathineni
DataForge is an AI-powered autonomous DataOps platform that detects data and pipeline problems, identifies their root causes, and helps automate their resolution.
insightforge-ai
Ansh Dhanwani
InsightForge AI is an autonomous business analytics agent that analyzes uploaded datasets to uncover trends, root causes, and actionable business insights.
Teen titan go
Nil Lad
Agent Guardian is an autonomous SRE incident-response agent that investigates production incidents by pulling real evidence — service metrics, error logs, and recent code changes — and proposes a root-cause fix, without ever being allowed to execute a destructive or unapproved…
Midnight
Prathmesh Adsod
Airstrong is airline disruption recovery software. It is built for situations where a problem that starts with one airport, aircraft, or flight quickly spreads across the rest of the operation.
Agent-Harness-Hackathon
Manishika Agarwal
FabGuard is an auditable maintenance-risk triage agent for semiconductor fabs and chemical plants. It reads structured maintenance events, calculates deterministic risk scores, generates reliability-focused recommendations, identifies evidence gaps and uncertainty, and prepares…
LocalRag
Nizamuddien TI
LocalRag is an AI career agent for job seekers. Instead of a chatbot that just answers questions about a CV, it actively works a job search: given a job description, it checks every requirement against the candidate's real CV — citing evidence for each skill rather than guessing…
applyguard-ai
Sakshi Tripathi
ApplyGuard is a human-in-the-loop AI agent that researches relevant career opportunities, identifies opportunities matching the user's interests, and presents them for review.
codegaurd
Tuhin Banerjee
A platform where you can check whether the codebase is safe for local installation or not just by putting the url.
DubSmash
Agastya Khati
ShutterFrame is a pre-production safety harness for database migrations. Today, teams often decide whether a migration is safe by reading SQL in a pull request.
MiloticX
Peyton Li
MiloticX takes a repo URL and walks through the README's setup steps one by one, running each command in an isolated sandbox to see what works and what doesn't. When a step fails it classifies the reason: a package got renamed, a flag was removed, or it's just a typo.
Byter
MAYANK MAHAUR
Byter is CI for bug reports. It turns incoming GitHub issues into deterministic proof, a tested candidate patch, and a maintainer-controlled draft pull request.
Endeavour
Kush Yadav
CHRONICLER is a persistent AI Dungeon Master for solo tabletop roleplay. Players speak or write a world into existence, choose an AI-generated quest thread, forge a character, roll D&D-style ability scores, and play through voice or natural-language actions.
resonance
Kshitij Jain
Resonance is for product managers who already have reviews and a dashboard that says “mostly positive.” Sentiment treats a polite 3-star “not bad” as okay; those customers often never file a ticket and then leave.
Ripple
Manan Paliwal
Ripple is a change-impact agent for supply chain and operations teams. When you need to rename a product SKU, Ripple reads your real operational database, runs deterministic impact math in a sandbox, shows you what is affected (purchase orders, shipments, customer orders…
tartarus
Yerramsetty Sai Venkata Suchita
1. Tartarus is an autonomous red-team and SecOps agent. It scans a GitHub repo, writes a real exploit, detonates it in an isolated sandbox to prove the bug is real, pauses for a human to approve, then opens a fix pull request. 2. Here the problem is scanners cry wolf.
aerosentinel
Jyotishka chattopadhyay
The project is primarily for organizations operating autonomous vehicles/robots where an AI can investigate incidents but should not be trusted to take dangerous actions without human oversight.
EraseGraph
Sahil Rakhaiya
EraseGraph is a privacy operations agent built around a simple idea: deleting personal data should end with evidence, not just a green “Done” message. When someone withdraws consent for a specific purpose such as using their data for AI model training—their information may…
kevinson
Spencer Jireh Cebrian
Cujo reviews pull requests by running them. A GitHub webhook starts one agent turn, the agent clones the PR into a throwaway Daytona sandbox, runs the test suite on base and head, writes and runs probes against the changed code, boots the app and hits it, and installs any newly…
Tracecause
Vibushan Vinayak K
TraceCause is an AI-powered causal investigation platform that helps users understand why a real-world outcome happened instead of just summarizing information. A user enters a company, event, or business outcome, and the system investigates it through multiple AI agents.
agent-harness-hackathon
Ashutosh Maurya
A retail store has 1,000 SKUs tracked in an Excel sheet. Instead of someone manually scanning it for low-stock items and calculating what to reorder, this agent does it — and pauses for human approval before writing anything.
Sudarshan
Ujjwal Mohan
Our project Connects 4 important tools to automate daily workflows safely, All with Human consent
querypilot
Akshaj Manchanda
QueryPilot is an evidence-first analytics and forecasting agent for revenue and operations teams at SaaS companies. It helps analysts investigate questions such as why qualified leads changed and what may happen next.
VSSP
Sabarish SK
OpsSentinel is an evidence-first incident response agent that investigates service failures autonomously but requires explicit human approval before taking destructive recovery actions.
pr-guard
Sridhar
PR Guard is a security-aware pull request review agent. You paste a GitHub PR URL. It reads the diff, runs the repo’s tests in an isolated sandbox, scans changed lines for security issues, drafts one review comment, then stops and waits for you before posting anything public.
ClanCode
Gaurav Chaudhary
ClanCode turns AI coding into a living game-like experience. Instead of watching an AI agent work inside a plain terminal, ClanCode represents your repository as a clan village.
cloudyyyy
james gurung
BlastShield is a safety layer between AI agents and production databases. It analyzes AI-generated SQL before execution, showing affected rows, cascade impacts, dependencies, and risk, then requires human approval before constrained execution.
portcullis
Arpan Patra
Portcullis audits an npm package before you let it into your repo, and it can't add it without your permission. The problem is that adding a dependency is the least careful thing most of us do.
PhoenixGen
Riddhiman Roy Chowdhury
Licence to Patch is an on-call agent that reads production errors from Sentry, investigates the code, reproduces the bug in a sandbox, creates a regression test, applies a minimal fix, runs the tests, and opens a pull request after human approval.
Ram And Co
Surya
AgentOpsML is an autonomous machine-learning experimentation system. It allows an AI agent to inspect a dataset, create reproducible train/validation/test splits, train multiple PyTorch model architectures, compare them using validation metrics, automatically select the best…
agent-harness
Abhijit Mohanty
OpenQuest is a human-in-the-loop agent for developers who want to contribute to open source but struggle to find the right repository, choose a suitable issue, or navigate the contribution process safely.
Orvexa
Toufiq Farhan
Orvexa is an autonomous database migration safety harness built for PostgreSQL engineering teams, platform engineers, and lead DBAs. The Problem: Modern CI/CD pipelines run migrations blindly: traditional runners consider a migration successful if PostgreSQL returns "exit code…
bountydesk
Vaibhav Tomar
BountyDesk is automated bug-bounty triage with a human sign-off. A vulnerability report arrives as a GitHub issue, an AI agent investigates it against a pinned target in an isolated sandbox, and the agent drafts a verdict with evidence: reproduced, not reproduced, or…
RECKON
Hillary Ikhais
RECKON is an agent built around one question: can I justify this action before I take it? It investigates real systems through MCP, builds an evidence-backed recovery plan, tests that recovery in a sandbox, red-teams its own reasoning, and then either asks a human to approve the…
ForgeCanary
Dhiran
ForgeCanary prevents MCP upgrades from silently breaking agent workflows. It replays successful jobs against both MCP versions, verifies the real external outcome, blocks regressions, approval-gates a scoped repair, reruns the jobs, and produces a release receipt.
UpgradePilot
JAIVIGNESH GK
UpgradePilot is a developer tool that verifies dependency upgrades before a pull request is created. It inspects a public GitHub npm/pnpm repository, shows outdated dependencies, runs the project’s real checks in an isolated sandbox, applies an upgrade, verifies it, and only…
codealongai
Krishna Kartik Darsipudi
Chat interfaces are linear, but learning, especially learning an unfamiliar codebase is not. Understanding code requires following relationships across functions, files, and concepts through multiple hops.
Spambots
Janesh Kapoor
SunoAI is a browser agent you drive in Hinglish that stops and asks before it ever writes anything. Ask it "analytics jobs dikhao" and it opens Chrome, navigates to the job listings, and reads the top results back to you, numbered so you can refer to one.
Managed
Tyra KJ
Airlink is an open-source remote control system for local AI coding agents running through True Forge. It lets developers start an agent on their workstation via a CLI or VS Code extension, pair it with a mobile phone using a six-digit PIN, send coding prompts remotely, view…
Warden
Piyush Raj
Warden is an incident response agent. When a production alert fires for a service, it investigates on its own: it pulls logs, metrics, and recent deploy history in parallel, then reproduces the suspected cause by actually running the candidate code in an isolated sandbox instead…
evidence-gated-runbook-executor
Sahil Silare
RunProof is an evidence-gated runbook executor for incident response. When an alert fires, its agent looks up the matching runbook, gathers only the evidence that runbook authorizes (logs, metrics, deploy history), runs a diagnostic script in a sandbox, assembles an evidence…
RailOps-Enterprise-DBMS
Shambhavi M K
RailOps Enterprise is a railway database management platform designed to centralize and simplify railway operational data. It provides an administrator-focused system for securely managing railway records through a structured, user-friendly interface.
omniforge
Jay Joshi
OmniForge is mission control for running autonomous agents against real infrastructure. It runs three specialized agents OpsForge for incident response, SecurForge for appsec, DataForge for data ops from one web cockpit.
ClipForge
krish
ClipForge is a TrueForge-powered video clipping agent. Upload a video, describe the moments you want in natural language, inspect and approve the agent's proposed tool call, and receive a rendered clip directly in the conversation.
sitrep-incident-commander
Nagarjuna
i have created SITREP, an incident responder agent. it analyses the server logs and other metrics by deploying sub agents. it suggest for change if it found any problem with the current deployment.
We-make-devs
sepal sagar
SafeOps AI is an approval-gated incident response agent designed for DevOps engineers and engineering teams. It autonomously investigates production incidents by analyzing monitoring metrics, application logs, recent deployments, and diagnostic results to identify the probable…
Null Pointers
Ronak Gupta
Autonomous, AST-aware, zero-downtime database migration & refactoring agent — built on TrueForge. SchemaForge treats a schema change as what it really is: a coordinated database and application-code change.
TypoBros
Aditya Kumar Puri
Researchers waste hours triaging papers: reading them, cross-checking claims, deciding what's trustworthy. Openwrite is an AI research agent that reads a paper for you and produces a Trail (what it did), a Coverage report, and a Claims↔Evidence table.
signalis
Ramkumar R
Signalis watches a company's CRM and website activity and tells a sales or marketing rep which leads are actually worth calling right now, and why. It reads in CRM rows and website events, figures out where each lead sits in the buying journey (early, mid, or late), explains…
trueforge
Sakshi Andhale
Trust Cop is an AI-powered security agent for MCP tools, built on TrueForge. It keeps a watchful eye on the tools your agents rely on tracking an approved baseline for each one and flagging anything that changes without warning.
AI_Repo_Analyzer
Aayush Bansal
AI Repo Analyzer is an autonomous multi-agent code auditing and technical due diligence engine. Manually vetting unfamiliar codebases takes 45–60 minutes per repo, while naive single-prompt LLMs suffer from high hallucination rates (34.5%) and missed vulnerabilities.
Agent-Harness-
Rajan mishra
Clearance is an agent for one job: a customer says they were charged twice. It looks up the order, proves the duplicate by running JavaScript in an isolated sandbox, then stops. process_refund is irreversible. Money does not move until a human stamps License.
aegis
Harley Albert C. Buendia
Aegis is an AI incident commander that investigates and fixes production incidents while keeping humans in control of high-impact actions. In our demo, a checkout service hits an 18.4% error rate.
flakebrake
Dan Vellon
FlakeBrake is a commitment firewall for humans and agents: it makes promise acceptance an admission decision instead of a model assertion, and it will not let an unsupported, stale, conflicting, duplicated, or unverified consequential claim cross an MCP boundary.
fuse
Dhyaneesh DS
Fuse is a governance and execution control plane for AI agents performing consequential work. Users visually author workflows that investigate an event, collect trusted evidence, rehearse an exact change, pause for human approval, execute through n8n, and independently verify…
We Crack Hacks
Nitin Pathak
VigilSRE is an autonomous AI Site Reliability Engineer (SRE) harness and incident response platform designed to eliminate production downtime. The Problem: When high-severity (P0) crashes occur (such as memory corruption, buffer overflows, or pool exhaustion), human engineers…
Barcelona
Krishna.R
Grid operators face incidents — storms, equipment failures, demand surges — where the "obvious" fix can cause a worse failure elsewhere on the network; GridMind's own hero scenario shows a proposed load transfer getting rejected in sandbox because it would have overheated a…
trueforge
Mahijith Menon
An agent that carries out GDPR "right to erasure" requests against a real production estate — and cannot destroy anything without a human saying yes.
00-Flake
Rohit
00-Flake is an autonomous AI agent that detects, reproduces, diagnoses, and safely quarantines flaky automated CI tests so they stop blocking release trains. The Problem: Every software team experiences intermittent test failures in Playwright, Cypress, PyTest, or Jest.
GoatWhistle
Mikhail
Harness is a 24/7 autonomous paper-trading agent for Alpaca. It monitors market prices, company news, macro movements and portfolio risk, then proposes explainable trades. The main problem it solves is safely turning an LLM’s probabilistic reasoning into real financial actions.
Redline
Rohan Jain
What Does Your Project Do? Redline Guardian is a TrueForge agent that provides safe, auditable SQL change management for databases. It takes SQL statements from operators and runs them through a 4-role safety workflow: Ledger — Analyzes the SQL to identify affected tables, rows…
ai-dungeon-master
ANUTTOMA SIL
An AI Dungeon Master (AI DM) is an AI-powered game master that runs a role-playing game by acting as the narrator, rule engine, and world controller instead of a human DM.
trueforge_android
Omkar Gothankar
TrueForge Android Operator is a governed computer-use agent for a actual Android devices. Speak or type a task on the phone - "dismiss all my notifications except WhatsApp" and the agent operates the phone through its accessibility tree, recovering with a vision sub-agent when…
cleanroom
Sreenath M Menon
Every team has one spreadsheet nobody wants to touch. Three people have edited it. Dates are written three different ways. There are duplicate rows from a double import, money stored as text with dollar signs and commas, and totals that don't add up.
DrHiro
Kresimir
gathers all private health data in one place and allows user to control and act on it
ModelAudit-Agent
Debabrata Pattnayak
A scan report per model with confidence-scored findings, not a black-box pass/fail; a registry-wide view showing every model's current certification status automatic re-scan triggers whenever a new fine-tune or compressed variant of an existing model appears
agent-harness-submission
Utkarsh tiwari
My project is an AI-driven Data Pipeline Assistant designed to automate exploratory data analysis. It solves the problem of manual, repetitive data manipulation by writing and executing Python scripts (using Pandas and NumPy) directly within a secure environment.
Code Crafters
Shashidhar Pawadashetti
IncidentPilot is an autonomous incident-response and postmortem agent for backend services. When a production alert comes in (e.g. "connection pool exhaustion on order-service"), it fetches real diagnostic logs and metrics, diagnoses the failure, reproduces the crash in an…
vextorn
Pranav Patil
Autonomous AI robotics agents often succeed in idealized simulations but fail catastrophically in the real world due to hardware realities: sensor latency, packet drops, CPU thermal throttling (85°C), and actuator lag.
countersign
Saharsh Tibrewala
When an AI agent wants to delete something from your database, every harness just shows you the raw SQL and asks allow or deny. You can't actually tell what's about to happen. How many rows die, what cascades, whether you can undo it.
gatekeeper
Anam khan
Gatekeeper — The Trust Layer for Autonomous AI Agents Gatekeeper is an approval-aware autonomous agent that gives AI systems bounded autonomy when operating real-world tools.
Avengers
Sri Nidhi Kuchana
Issue Resolver Agent automates the manual loop of resolving a GitHub issue: reading it, digging through the codebase to find relevant files, diagnosing the bug, writing a fix, and opening a PR.
Runtime Rebels
Eva Chen
RootCheck is an AI agent security testing harness that evaluates how agents behave under adversarial conditions. It runs controlled security scenarios, such as indirect prompt injection, observes the target agent’s actual tool calls, and determines whether the agent followed…
drawn
Mintu Gogoi
An agent connected to an MCP server normally answers with a wall of JSON or a paragraph describing what it found. drawn makes it answer with interface instead. Ask for flights and you get fares you can click.
exitramp
Siddharth Choure
ExitRamp is a pre-migration gate for engineering teams replacing the model behind a tool-using AI agent. A final-answer-only evaluation can miss a dangerous regression: a cheaper model may sound correct while skipping required tools, using the wrong arguments, or making claims…
BlackBox
Aayushi Gupta
Sentinel is an autonomous crisis intelligence and emergency dispatch platform built for 911 dispatchers, emergency operations centers (EOCs), and first response commanders who face severe sensory overload and coordination delays during chaotic disasters.
rocky-code
Aditya Sasidhar
I built Rocky, a terminal coding agent for developers who want the speed of AI coding tools without giving them unrestricted access to a real working tree. Rocky delegates work to Codex, Claude Code, or OpenCode inside disposable containers.
ScholarAgent-Autonomous-Research-Reproduction-Agent
Tanya Garg
ScholarAgent is an autonomous AI research reproduction agent that helps researchers, students, and ML engineers verify whether the experimental results reported in a research paper can actually be reproduced.
trueforge-ai-buyer
Tejas M
Will mail it soon
trueforge-rote-fix
Albert Joshua
This is a general deep research agent that can deep search the web for content using sub agent spawning and produce analysis reports
scope-city
Rajdeep Singh
Give an AI agent a Stripe key so it can refund one charge, and it can refund all of them. The same is true of GitHub, Salesforce, or your database: every integration hands over a credential that can do far more than the job needs, and afterwards nobody can say what the agent was…
The Farmers
Shanu Kumawat
TrueStrike is an autonomous pentest agent that asks permission. Point it at a target you own and it recons the whole attack surface, proves every finding with a working exploit inside an isolated cloud sandbox, then stops and waits for you before anything destructive.
WeMakeRevenue
Rudra Veer Singh Rathore
Recoup is an SRE-style investigation agent for revenue, built on TrueForge. Point it at a failed-payment spike, a refund wave, or a renewal-risk signal, and it works the incident the way an SRE works an outage: pull evidence from Stripe, Sentry, GitHub, and tenant data via…
greenlight
Aryan Kumar
Greenlight is a CapCut-style cloud video editor with an AI Producer beside the timeline. Creators can drag, trim, split, mute, add transitions, mix audio, and undo normally, or ask the Producer to research, script, find licensed media, edit, render, and stage to YouTube.
vendor-scout
Shivam Arora
Vendor Scout is an autonomous procurement agent for hardware teams. You give it a part you need, your current supplier, quantity, price target, lead-time requirements, and technical constraints.
PodNot
Akshat Arya
ReDessIo is an autonomous spatial interior architecture and procurement agent harness built on the philosophy: "Watch it design. Approve before it spends." Problem It Solves: Generative image AI hallucinates beautiful interior concepts that can neither be purchased nor fit real…
sentinel-008
Dhruv patel
Sentinel·008 is an approval-gated incident response agent.When a payment-failure alert fires, it investigates the incident the way a senior SRE would: it reads dashboards, queries the database, searches logs, examines recent deploys, and runs diagnostic code inside an isolated…
operationsos
Shivam Srivastava
OperationsOS is an Business Operations Management Solution for a company run by AI agents a walkable 1-bit office where user is the CEO. Five specialists sit at desks doing real operational work on a live business: a Data Analyst querying the customer database, Market Research…
Blast Radius
Nishad Mulay
Blast Radius is a safe AI agent for deleting a customer’s personal data from a database. It first finds every record connected to that customer, such as addresses, uploads, support tickets, orders, audit logs, and order items.
Shaken, Not Staged
FNU Smriti
Production incidents demand fast decisions, but rushed remediation can worsen an outage. On-call engineers often need to manually gather metrics, logs, traces, deployment history, and ticket context before they can determine whether a rollback is safe.
agent-harness-automation
Payas Dhone
An AI agent system running on TrueForge that fetches real-time repository trends, executes analytical Python code inside an isolated Daytona sandbox, and enforces human-in-the-loop approvals before file operations
Warden
Naman Jain
Warden is an autonomous release gatekeeper built on TrueForge. It reviews GitHub pull requests by fetching real PR data through GitHub MCP, delegating security, style, and testing analysis to specialist subagents, running the PR’s exact commit inside an isolated Daytona sandbox…
aegis
Abhishek Patel
AEGIS is an autonomous runtime reliability agent for Node.js. It finds event-loop starvation (a CPU-bound task — massive JSON parsing, regex evaluation, sync crypto — blocking the single thread), reproduces the exact failure under a controlled load experiment, applies a targeted…
stamp
Fawuzan
Vendor reminders can look identical to new invoices, creating a real risk of paying twice for the same service. Using TrueForge, we built an AI agent that checks the books before drafting a payment-confirmation message.
Polaris-CLI
Dilip Kumar R
Polaris AI runs an ensemble of agents to coded implementations any Machine Learning paper. It aims to bring the knowledge of using coding agents to ML papers, to cut down on time taken for the implementations but instead focussing on running the experiements on our own platform…
docxy
Arindam Majumder
Docxy is an automated documentation and changelog pipeline powered by five specialized agents and human approval. Every push to main triggers agents that analyze code changes, map their impact on docs, update the docs, write changelog entries, and coordinate the final result.
trading-desk-sentinel
N S Siddarth
Sentinel Trading Desk is an AI-powered incident response and risk-monitoring system for paper-trading accounts. It connects to Alpaca through an MCP server and uses an AI agent to investigate unusual portfolio events by checking account state, portfolio history, recent orders…
Innovators
Kamal
SafeRun is a database guardian agent. In July 2025 an AI coding agent deleted a production database, ignored an instruction to stop, and claimed the data was unrecoverable. SafeRun is the layer that was missing.
picto
Sahil Gupta
I have noticed a problem of overwork of maintainers of open-source organisations. Due to AI development they are facing many sloppy prs/issues.
colophon-agent-harness
Manann arora
Colophon is a spec-first product video-generation harness with taste. An agent writes a video plan, Colophon enforces a closed motion vocabulary and runs timing, text overflow, brand-colour compliance, source-backed claims, and more.
fuseops-trueforge
amit mishra
FuseOps is an approval-gated incident-response agent for SRE and platform teams. When checkout payment failures spike, it gathers incident context, service health, deployment history, sanitized logs, and the runbook through an owned MCP control plane.
YukClara
P Koti Darshan
SENTINEL is an autonomous incident-response agent for Kubernetes. When a Prometheus alert fires (CrashLoopBackOff, API error spike, DB timeout), it loads the matching runbook, sends four sub-agents to investigate in parallel, the K8s state, logs, API metrics, and DB metrics over…
prism tech
Dheeraj kumar
When a bug/error shows up in production, this agent investigates it on its own, finds the root cause, proposes a fix, and verifies it in a sandbox — all without a human manually digging through logs and code, needing only one approval at the end
falcon-harness
Saksham Mishra
Falcon is a diff-scoped exploitation agent. AI coding agents ship new endpoints fast, and every so often one forgets an access-control check — and traditional scanners only guess a severity.
memoforge
Aditya Gupta
An analyst at a boutique venture advisory or VC firm receives a founder's pitch — a deck, an email, a set of notes — and has to turn it into a clean, consistent, house-format Investment Committee (IC) one-pager that a partner will read, trust, and act on.
clinical-ops
Sakeena Patel
This project is to automate healthcare administration reviews.
vetit
Victory Lucky
MCP servers are becoming more popular, in the same way it is being met with security issues and vulnerabilities, and according to AgentSeal 2026 report, 66% of 1,808 MCP servers had a security finding, that's roughly 1,193 compromised servers.
NOTEAM
Arav Latiyan
An agent that takes a suspicious email, detonates its links in a sandbox, and runs three parallel investigations (infrastructure, identity, history) into a verdict.
migration-harness
Yashkumar Kavaiya
Autonomously modernizes a bounded .NET service to Rust, then proves behavior before anyone can cut over. Compilation is not enough. Teams ship migrations that compile and still break production. This is for teams who need a cutover they can defend.
Null Garden
Bhavay Goel
The problem Webhook systems can tell you that a delivery failed. They cannot necessarily tell you whether the business operation failed. The provider records a failed delivery, but the business mutation has already occurred. Blind replay can perform it again.
forge-guard
Shoumik Chandra
ForgeGuard is an enterprise-grade AppSec governance layer designed for security engineers and platform teams deploying AI agents. It solves a critical vulnerability known as the "Lethal Trifecta" when an agent possesses simultaneous access to untrusted input, private data, and…
debugforge
Priyanshu
DebugForge is an autonomous debugging and code-repair system for developers. It tackles the growing difficulty of debugging complex and AI-generated code by reproducing failures, tracing likely root causes, generating surgical fixes, and verifying those fixes before applying…
graft
Md Abid Hussain
Graft is an agent memory for hackathon stacks. Every WeMakeDevs hackathon runs on a different stack, and the context an agent builds reading its docs dies with the session — so every event starts with re-reading the same documentation.
quartermaster
Manu Mishra
Coding agents claim. They say "I've fixed the failing test" and you find out in CI, twenty minutes later, that nothing ever ran. The claim is cheap because nothing forced the agent to execute anything. During development this agent did exactly that.
LUPIN-007
Ramanpreet singh
LUPIN is an agentic SRE assistant that replaces panic-driven guesswork with safe, reproducible diagnosis. The problem: When an alert fires at 3am, an SRE either blindly SSHs into production (risking making it worse) or sits paralyzed.
agentguard
Jaival Suthar
AgentGuard is a runtime trust layer for AI agents. I built it to control what an agent is allowed to do and independently verify what actually happened during execution.
taro
Chetan Patil
Today's AI agents talk to one person and work for one person. Real work isn't like that , what if we want to talk to the different people to schedule , coordinate to make a conclusion of a job , for example a single roofing repair means coordinating a homeowner, a project…
confirm-deny
Anik Das
CONFIRM/DENY turns an unverified bug report into a verified case file, or a defensible “cannot reproduce” with proof. Give it a GitHub issue URL. It runs the reported code in a credential-free sandbox, proves what ran, bisects to the first bad version, builds a structured case…
approvedeck
Aarav
Agent harnesses pause before anything irreversible, and that pause is the whole safety story. Today it is buried in a chat transcript. Run five agents and the approval that matters is three tabs deep and forty messages up, so people either miss it or rubber-stamp it later…
ROOK
Jay Rietzke
ROOK is an evidence-first AI incident commander built on TrueForge. It investigates incidents using least-privilege, read-only tools, correlates TrueForge tool calls and responses into retained evidence, and only promotes operational state when that evidence satisfies a strict…
rag-doctor
Kaushal Francis J
Checks current RAG system health and if rag is bad, runs diagnosis, and fixes.
repoguard
1
RepoGuard is an autonomous health and repair repository agent that inspects repository code and CI results and identifies the root cause of failing test cases, diagnose, make minimal repair and requires human approval before any repository change.
revenue-sre
Siddharth Jha
Revenue SRE is an AI-powered payment incident commander for online merchants using Razorpay. Payment dashboards show individual failures, but merchants can struggle to recognize when many failures represent one larger incident.
trueforge-proofboard
Manuele
Proof Board is a control plane (in Kanban style) for autonomous software delivery. It lets AI agents implement changes, independently reviews the actual sandbox artifact, and requires explicit human approval before publishing a pull request.
mohenjodaro
Aryan Singh
It helps in building testable evals for any robot on environements based on requirements and run eval sets in isolated sandboxes and relay the results in a research scoped UI desined for robotics engineers and researchers.
ResearchForge-AI
Girivasan B
ResearchForge AI is a multi-agent autonomous research platform that transforms a single query into a verified, citation-backed intelligence report. The system uses specialized AI agents, MCP tools, and sandbox execution to gather, analyze, validate, and synthesize information…
GuardForge
Aman Singh
Problem: AI agents given access to codebases have no safety layer — they can merge broken code, run untested fixes, or deploy to production without human oversight.
PATCHPILOT
Swayam Prakash Sahu
PatchPilot is an AI production engineer, built on the TrueForge agent harness — it automates the tedious middle of incident response while keeping a human in charge of the one irreversible step.
hello-agent
Divyansh
hello-agent (Repo Maintainer) is an agent that scans a GitHub repository for real, verifiable code issues (lint and static-analysis findings — not guesses), proposes concrete fixes, and opens a pull request once a human approves the changes.
fix-radar
Jayesh Agarwal
Fix Radar closes the loop between code review and code fixing. Qodo reviews a pull request and flags a real issue; a TrueForge agent picks up that finding, reproduces it in an isolated sandbox with its own regression test, writes the smallest correct fix, verifies it, and opens…
Rootless
Onkar Tulshigar
Domain Hunter finds suspicious domains that may be impersonating a brand. It checks the domains, looks at current and historical evidence, and gives a simple risk assessment for a human to review.
ripcord
Aman
Ripcord is a simulation-first, on-chain incident response agent for DeFi protocols. In the event of a live exploit (like a reentrancy attack draining a vault), speed is critical but autonomous agents are dangerous if given unchecked write access to production smart contracts.
thExplorers
Saurabh Gupta
Wrong fixes. Conversion drops, three dashboards show the same number, a PR goes out, and the real cause was a deploy or a log line nobody pulled. LOOP is for the person who owns the product.
DevGuard
Aniket Patel
DevGuard is a control plane for agentic software development. It lets engineering teams run AI agents that can inspect repositories, modify code, execute changes in an isolated sandbox, open pull requests, and respond to review feedback - all under explicit, enforceable policies…
harness-os
Harsha Priya Ganapathy
Harness OS is an adversarial crash-test lab for autonomous AI agents. As AI agents move from answering questions to taking real-world actions—issuing refunds, modifying repositories, deploying software, or calling external APIs—they can fail in ways traditional software tests…
Sentinel
Mayank Sharma
Sentinel is an autonomous AI security investigator that correlates identity and network evidence, produces an auditable threat assessment, and proposes containment.
coercion-brake-trueforge
Vanshika Jadam
Project Name: Asset-Hostage Shield: The Digital Coercion Brake The Problem: Ransomware attacks and scams weaponize extreme urgency (e.g., a 1-hour countdown to leak assets or pay Bitcoin). This induces panic, driving users or security teams into hasty, unverified decisions.
patchforge
Vivek Silimkhan
Developers who maintain dependencies. When a vulnerability is found, you need to upgrade, refactor the broken code, test it, and approve it, that's a day's work every time.
sentinel-agent
Prince Panchani
sentinel-agent is an autonomous incident responder for production software, built on the TrueForge harness. Give it a production incident and it investigates end to end — reads the incident, characterises the symptom, enumerates recent deployments, reads the actual diffs…
Offline
Offline Project Submissions
Handed in live in the room in San Francisco, on August 29, 2026.
55 projects
agentlens
Weigang Geng
AgentLens is observability for AI agents. Teams running agents on TrueForge have no way to see what their agents are actually doing - which sessions failed, which tools are flaky, where tokens and money go.
Aloha Live
Seth Caldwell
Aloha Agent connects travelers, locals, and the nonprofits serving the places they visit. Through voluntourism and Aloha Live experiences, our agent continuously researches what’s happening in a community, ingests live signals from travelers, verified locals, nonprofits, and the…
itinerary-agent
John Huang
Compass is a agentic travel planner. I have traveled a lot, and it always takes hours to research and make the plan. This is my attempt to create an agentic app for consumers.
DevGuard
Aniket Patel
DevGuard is a control plane for agentic software development. It lets engineering teams run AI agents that can inspect repositories, modify code, execute changes in an isolated sandbox, open pull requests, and respond to review feedback - all under explicit, enforceable policies…
hometown-grand-prix
Karthik Shetty
Street Steward is a 3D racing game on real OpenStreetMap roads (Downtown San Jose) where AI agents govern the race instead of hard-coded rules. An AI Traffic Steward watches your speed against the real road's speed limit and rules on violations in real time — but the driver must…
DreamTeam
Gracelyn Newhouse
Agents are overconfident about method. Given an unfamiliar task, they improvise a plausible approach instead of consulting the published research that already solved it, and then report success on checks that never ran.
aug29
Tri Nguyen
It's for on-call engineers and SREs. The real pain isn't fixing the bug — it's the first ten minutes of an incident spent figuring out what even happened and whether it's your fault or a vendor's.
parity-agent
Bryan Samuel James
parity-agent detects personalized and dynamic pricing, the kind you can never spot yourself since you only ever see your own price. It fetches the same product page from isolated sessions in different countries and compares what each one is shown, distinguishing real…
Agent Fight Club
MOHAMMED ADNAN
It's a 3D city where every building is a real GitHub repo. Click one, pick an issue, and two AI agents race to fix it in sandboxed clones while their commits stack up as tower floors, then a referee re-runs the original tests, code quality scores both diffs live, opens the…
agent-harness-hackathon
Akshata Madavi
Problem: AI evaluation platforms show grader scores and traces, but when a score is low, engineers still have to manually inspect traces, grader logic, prompts, tools, and code to understand why it failed and whether the agent or the evaluator is actually at fault.
beforeyoupay
Lea Frances Yu
BeforeYouPay is an AI accounts-payable investigator that checks suspicious invoices before a business sends payment. It is designed for accounts-payable teams, finance leaders, operations staff and small businesses that may not have dedicated fraud analysts.
Team JC
Carl Okpala
We are shipping web software faster than ever, and almost nobody checks whether disabled people can use it. The tools that exist read the code but never use the site. A button can say "closed" in the code, open a menu when you click it, and still say "closed" afterwards.
gofer-trace-smb
Jason Okorie
Small companies struggle to harness AI agents when it comes to company-specific context. Gofer Trace helps companies give agent harness the context needed to pilot agents across different tasks. Local businesses setup LLM knowlege bases to help give insight on day-to-day tasks
PharmaFlow
Ramya Velaga
PharmaFlow helps pharmacy and healthcare teams keep track of FDA safety alerts, drug shortages, inventory, and medication availability in one place. Its agents collect live information, work together, and notify teams when something needs attention.
Ratchet
Shryuk Grandhi
Every step is a git commit plus a sandbox snapshot, and every candidate patch must clear a seven-stage verifier gauntlet (build, cheat check, fail-to-pass, pass-to-pass, types, lint, diff hygiene) before it sticks.
Cross Exam
Martin H. Pefaur
CROSS-EXAM is a dry run for AI actions that cannot be undone. When an agent proposes an irreversible action — a bulk refund, a payout, closing an account, every guardrail on the market grades the action the agent declared.
petassist
Yu Iwase
Petassist makes AI agent permission visible to people who are not engineers. Pets live in a pocket window and stick to the monitor like notes. A cat investigates untrusted mail inside TrueForge’s sandbox and cannot send. A bunny drafts a refusal and cannot send.
SOLO: Aishow "詠唱" — "incantation" in Japanese.
Yoshihiro Sekiya
Aishow ("詠唱" — incantation) is a voice-summoned AI agent for macOS. The problem: as a Japanese engineer at a US startup, every cold outreach through a company's website contact form costs me ~15 minutes — research the company, draft in English, fix the translation, paste it in.
ember
Sviatoslav Zinevich
I decided on creating an incident resolver "Ember": Submit your critical software error and the agent will identify the error, explain it, and resolve it. From pool exhaustion, worker/thread starvation to poisoned-config, poison-message, stuck queue, and more.
Meeting-Action-Agents-
Vidya Jayaraman
What does your project do? Meeting Action Agent turns messy meeting transcripts into actionable next steps. It uses a multi-agent workflow to extract decisions, action items, owners, deadlines, and open questions, researches the company discussed using Bright Data, and generates…
data-harnesses
Thomas Barrios
made just for me to track my macros and health
Hush-Harness
Shradha Pujari
Hush is an autonomous data-center incident triager and operator. It turns a large alert storm into one evidence-backed root cause, enriches the incident across infrastructure systems, proposes remediation, executes safe actions, and requires a human decision before any…
fde-robot-soc-harness
suhaas
Physical SOC turns a Reachy Mini desk robot into the human-in-the-loop approval gate for automated vulnerability remediation. Scanners that find CVEs and bots that write patches are solved problems.
ForgeRoom
Pramod Thebe
ForgeRoom is a shared workspace where people collaborate with persistent AI coworkers—not one-off chatbots, but durable teammates with their own sessions, permissions, and audit trail. Problem: Most AI agent workflows are opaque and hard to trust.
countermove
Sterling Cobb
Countermove helps you make any sort of business decision by simulating outcomes
yallapost
Taha
YallaPost is a daily content agent for creators on the startup and venture beat. The hard daily decision is what to make today, so the agent scouts a watchlist of Instagram, X, search, and RSS sources, surfaces three trending topics with clickable evidence, writes a beat-by-beat…
ai-video-studio
AI HOLLYWOOD
Touchdown AI Video Studio turns a plain-English creative brief into a structured story, six-shot storyboard, production prompts, and an approval-controlled video render plan.
mcp-sentinel
Kiran Devihosur
mcp-sentinel checks whether an MCP server is safe before your team trusts it. It reads the server's tools, runs the server in a sandbox with fake secret values, calls its tools, and watches whether it tries to steal those secrets or hides instructions in its tool descriptions.
Winning Team
OUSSAMA ELFIGHA
UpgradePilot turns a dependency-breaking change into a verified pull request. A developer provides a GitHub repository and upgrade target. UpgradePilot retrieves live migration documentation, identifies affected call sites, reproduces the breakage in an isolated sandbox…
RoolyTooly
Raphael Khalid
- roolytooly is a self-learning harness: an agent that stops repeating its own mistakes. - Coding agents fail in the same families of ways over and over: reporting a stale artifact as fresh, calling a suite green after running one file, claiming "done" on an empty export…
Everest G1
Kaushik Sivakumar
Everest G1 rescue ops are difficult and takes a lot of time in extreme mountain terrain, where reaching an injured or lost climber can put additional human rescuers at risk.
truthlease
Rahul Krishna
TruthLease is an agent safety harness that helps operators safely fix previously valid actions when conditions change. It monitors action contexts for stale evidence, re-validates the facts, routes results through a transparent analysis and approval flow, and applies only…
whisperer
chris mckenzie
It finds out why people are switching to or leaving your product, discovers difficulties they are having, confirm if the bug or issue is real, patches it if it is, follows up with the user and improves the product.
Sponde (team name), SOLO
Abhinav Garg
Sponde is a deal room for AI agents. Everyone is about to have a personal agent, but there is nowhere safe for two of them to make a deal. In Sponde your agent and my agent negotiate through a neutral room where only validated offers can cross, so private stuff like salary…
reweave
NarasingaMoorthy V
Every team that extracts data from the web is on the same treadmill: r/webscraping practitioners report 10–15% of scrapers breaking every single week as sites change structure, and one person maxes out maintaining ~100 of them.
artiji-buyer-agent
Robert Schwentker
The project solves the trust & reliability gap in paid, deferred MCP services. It enables an AI buyer to inspect material terms before payment, obtain human approval, pay exactly once, survive restarts, and verify that the eventual artifact matches the purchase.
mcp-breaker
Dylan Flud
MCP Breaker is a security testing tool for AI agents that use MCP. Giving an agent access to tools like GitHub, files, databases, or internal services is easy.
TEAM BRAHMA
Sai Nellutla
BRAHMA is a fault-tolerant command center for autonomous AI workers. It takes a production incident from alert to verified remediation: investigating the issue through MCP tools, dynamically creating specialized AI workers, recovering when workers fail, testing AI-generated…
Edit-AI
Sunny Rodrigues
EditAI is an open-source video editing agent. You upload a video, tell it what you want in plain English, and it plans the edit, writes its own ffmpeg commands, and runs them in an isolated sandbox.
OPESHA Credit Intelligence Agent
John Wangombe
OPESHA Credit Intelligence Agent is an agentic AI system designed to help lenders assess SME credit readiness faster and more intelligently. The agent researches an SME, gathers relevant publicly available alternative-data signals, analyzes supplied financial information…
MineForge
Andre de la Cruz
MineForge lets AI agents operate as embodied workers inside Minecraft, where they can gather resources, hunt, craft, and construct imported blueprints. It gives Minecraft players and agent developers a safe way to experiment with autonomous, observable multi-agent behavior.
Incident Responder
Mayuresh More
Ripcord is an AI on-call engineer. When a production alert fires it investigates with read-only tools, writes and runs its own diagnostic code in a sandbox to quantify what it found, proposes exactly one remediation — then stops dead and asks a human before doing anything…
warren
Kian Kyars
Ingests a ChatGPT export, clusters unfinished threads, researches the stale ones against a local corpus, drafts a morning plan, and does not write that plan until a human commits it.
scholarly
Pulkit Arya
Scholarly
hold-the-line
Karim Baba
Hold the Line is a phone number you can call. On the other end is a claims agent for a fictional insurer working a vehicle total-loss claim: it pulls the policy, values the car, and computes the settlement. The point is the console beside it.
dgx_hack
Vaibhav Maheshwari
Agents take a lot of time to do their work and which is why people just run them and let it YOLO stuff, this is great but it creates a problem where the person who gave the assignment for the same does not have the visual feedback for the work done, what we have done is created…
Dgx
ali amjad
WaterOps is a role-based AI operating layer for industrial water treatment plants. It connects specialized AI agents to an existing Ignition HMI through MCP, with TrueForge acting as the local agent harness.
upstream-watch
Kush Ise
Upstream Watch reads the things your code depends on — hosted APIs like OpenAI and Stripe, and open-source packages like Express and React — and finds the changes that will break you before your users do, then patches the code, opens a pull request, and stops and waits for a…
Nolan
Aditya Bhat
Engineering ships faster than marketing can distribute. That's the gap Nolan closes. Product and engineering teams now ship weekly, sometimes daily. Growth and marketing can't produce launch content at that rate: every feature wants a demo video, a LinkedIn post, an X post…
agent-hackathon
Suman Polepaka
LearnLoop — From content to capability. LearnLoop is a visual, persistent learning agent that turns confusion into a personalized 8–15 minute practice loop: See → Predict → Practice → Prove → Advance Learners can paste a YouTube link, describe where they are stuck, or submit…
Team Wingman
Smaran Aramballi Sandarsh
Wingman is an AI agent that handles the early stages of dating for you, and stops to ask permission before it says anything you did not authorise. You never fill in a form.
placebo
Hari Gangadharan
I am a game dev on UGC platforms like Roblox & UEFN, but I face an issue when using AI- it cannot debug games (even with tests, it does not test if it actually made a difference, since the models reward hack!), and I cannot run it locally: - I made Placebo to solve this, since…
trueforge-agent-harness-hackathon
David Van Gheel
It is for people looking for jobs. The AI will prompt the user for information about what their skills are, what they are looking for an other relevant information. Then it will give you a list of jobs to apply for and help you to adapt your resume to fit the jobs.
BizForge
Yaroslav Volovich
BizForge is agentic system which helps entrepreneurs find business or product opportunities. 1) First "setup" agent interviews the user (could be just LinkedIn link) collecting user competencies and business preferences (eg avoid military).
two-fold
Shubham Shinde
Twofold is a website audit for the two audiences every public site already has: people and AI agents. Today, teams measure humans with analytics, heatmaps, and A/B tests. Agent visitors — copilots, research agents, scrapers — hit the same pages, fail silently, and leave no score.