GitHub projects I’ve starred, and what each one adds to the state of the art. Repos are grouped into themes, and every review links to the repos it builds on or competes with.
GitHub projects I’ve starred, and what each one adds to the state of the art. Repos are grouped into themes, and every review links to the repos it builds on or competes with.
56 items under this folder.
Each theme is a running summary of one area: the main approaches, what each repo added, and the gaps. It's rewritten whenever a new repo joins.
Installable skill packs and plugins that extend what a coding agent knows how to do, from full development workflows to domain-specific judgment. Approaches range from composable single-task skills to fixed methodology sequences to structured reasoning frameworks.
Main approaches: Skill collections, Domain skills, Development methodologies, Structured reasoning lenses
8 repos · updated Oct 10, 2026 · Read the theme →
Workspaces, IDEs, and platforms where agents and people coordinate work together. Approaches split between desktop IDEs that wrap CLI agents, shared event-log workspaces, self-hosted session daemons, and AWS-native control planes.
Main approaches: Agent IDEs, Shared workspaces, Platforms on AWS
7 repos · updated Oct 10, 2026 · Read the theme →
Languages, auditing tools, creative-coding libraries, and personal utilities for people who make software and things. Approaches range from compile-time language design to browser-based creative sketching, web auditing, and window management.
Main approaches: Languages, Web auditing, Creative coding, Personal utilities, Hardware
6 repos · updated Oct 10, 2026 · Read the theme →
Benchmarks, leaderboards and research environments that measure how well agents or models do defined work. Approaches range from live-cloud and stateful-mock task suites to harness comparison studies, decision-model scoring, and multi-agent simulation.
Main approaches: Harness benchmarks, Domain task benchmarks, Decision-model scoring, Research environments
5 repos · updated Oct 10, 2026 · Read the theme →
Libraries, stores and models that give agents memory and durable context across sessions. Approaches range from plain-text prompt scripts to cloud storage backends to research models that rewrite their own context.
Main approaches: Memory prompts and scripts, Storage backends, Context databases, Context models, Session trajectories
5 repos · updated Oct 10, 2026 · Read the theme →
Policy languages, sandboxes, and security scanners that limit or audit what an AI agent may do. Approaches range from static analysis orchestration and structured threat modeling to a Cedar-derived policy language with temporal reasoning and OS-level sandboxing.
Main approaches: Policy languages, Sandboxes, Security review tools
5 repos · updated Oct 10, 2026 · Read the theme →
Databases, pipelines, encodings, and generators that store, move, or produce data. Approaches range from binary wire formats and distributed key-value stores to bulk ELT connectors, live event subscribers, and synthetic document generators.
Main approaches: Databases, Data movement, Encodings, Synthetic data
5 repos · updated Oct 10, 2026 · Read the theme →
Tools for running models on your own hardware or in your own cloud account. Approaches range from multi-cloud compute schedulers to bare-metal serving control planes, local desktop apps, and cloud-provider sample notebooks.
Main approaches: Compute schedulers, Model serving, Local AI apps, Cloud model samples
5 repos · updated Oct 10, 2026 · Read the theme →
MCP servers, CLIs, and discovery specs that give an agent a callable capability or let it find one across networks. Approaches split between domain-specific tool wrappers and the plumbing that connects agents to those tools.
Main approaches: MCP servers, Discovery protocols
4 repos · updated Oct 10, 2026 · Read the theme →
Small models that make one typed decision or classification instead of calling a full LLM. Approaches split between task-specific classifiers, general-purpose decision models, and ecosystem indexes around a single model.
Main approaches: Decision models, Ecosystems, Task-specific classifiers
3 repos · updated Oct 10, 2026 · Read the theme →
Frameworks and formal work for building a single agent: its loop, tools, state and runtime. Approaches split between runnable harness implementations and composition theory.
Main approaches: Harness frameworks, Composition research
2 repos · updated Oct 10, 2026 · Read the theme →
Repos in this theme
55 repos
What's new: Goose eliminates the heap entirely: all dynamic memory lives on bump-pointer stacks, so freeing any structure costs one store and references into growing arrays stay valid without lifetime annotations.
The Goose Programming Language
What's new: PixlPut tells windows of the same app apart by workspace path or document, and by open browser tab once you turn on an opt-in setting, so each window can return to its own saved spot.
What's new: Tasks target multi-step business workflows on local replicas of real commerce platforms, and each task passes only when every per-task verifier check clears against actual state changes in an isolated mock service.
CommerceAgentBench: Benchmarking Long-Horizon Agents in High-Fidelity, Stateful, and Reproducible Replicas of Real Online Services
What's new: Airbyte pairs its ELT connector catalog with a separate Agent SDK that turns connectors into typed LLM tools for pydantic-ai, LangChain, OpenAI Agents, and FastMCP.
Open-source data movement for ELT pipelines and AI agents — from APIs, databases & files to warehouses, lakes, and AI applications. Both self-hosted and Cloud.
What's new: OpenResearch adds literature grounding against alphaXiv, bioRxiv, and PubMed so agents can form informed hypotheses, and records every experiment in a local SQL database so agents can read past results before deciding what to try next.
Turn your coding agents into research agents
What's new: ARD defines a domain-anchored, federated discovery layer so agents can find callable resources across networks of discovery services.
Agentic Resource Discovery (ARD) specification
What's new: It maps the SDK's slash-separated key contract directly onto DynamoDB partition-plus-sort keys, with optional vector search, TTL expiry, and S3 offload for oversized values and no separate vector store.
Amazon DynamoDB storage backend for the Strands Agents SDK (TypeScript and Python): sessions, agent memory, semantic search over DynamoDB vector indexes, and S3 offload behind one Storage interface.
What's new: aws-bench provisions real, isolated AWS accounts, deploys CDK stacks into them, runs the agent with scoped credentials inside a container, and then scores the result with an LLM judge or a check of actual AWS resource state.
aws-bench measures how well AI agents and model combinations perform on real AWS work — diagnosing misconfigurations, provisioning infrastructure, and operating live cloud environments.
What's new: This repo collects AWS-hosted examples specifically for Anthropic Claude, covering prompt engineering, tool use with complex schemas, PDF knowledge bases with citations, Intercom classification and a multimodal Streamlit playground.
What's new: It offers two concrete compute paths, EKS with Helm or ECS plus Bedrock AgentCore Runtime, sharing the same agent interface and gateway layer, so teams can evaluate operational trade-offs against a working codebase.
A sample agentic ai platform to run agentic workflows on AWS using either EKS or Bedrock AgentCore with open source frameworks like LangChain/LangGraph, Strands, etc..
What's new: AVA covers the full lifecycle from business-case DCF models and operating-model selection through to compliance checklists for 14 frameworks, all surfaced in a single React/FastAPI control plane deployed on ECS Fargate.
A platfrom to Plan, Build, Operate, Secure, and Govern AI Systems for Financial Services on AWS
What's new: This collection has one Kimi K3 notebook for Amazon Bedrock and links to the Kimi API Platform on AWS Marketplace and to external SageMaker deployment examples.
Samples for getting started with Kimi models by Moonshot AI, on AWS
What's new: It treats document structure and tabular data as separate concerns: PDF-to-synthetic-HTML conversion runs through Bedrock Claude, while a four-stage pipeline (extract, schema, generate, test) splits paid schema extraction from free, offline, seeded row generation.
What's new: Rather than wrapping a single scanner, ASH orchestrates a roster of built-in security tools, normalizes all findings into SARIF/JSON/HTML, and exposes the whole pipeline as an MCP server so AI coding agents can trigger scans and read results directly.
ASH is an extensible, open source SAST, SCA, and IaC security scanner orchestration engine.
What's new: It structures threat modeling into nine phase-gated steps, from business context through residual risk and export, plus an optional code-validation step, with completion checks before export.
What's new: Buzz puts agents and humans in the same signed event log, using Nostr keypairs for identity, so every message, patch, CI result, and workflow step lands in one searchable audit trail rather than across separate tools.
A hive mind communication platform
What's new: Tardigrade derives durable agent state from an immutable event log through composable atoms, drawing on inspiration from Elm.
The TypeScript framework for building modular agents around an immutable event log.
What's new: Tastemaker uses runnable Python scripts (contrast checker, palette generator, anti-slop scanner) and writes decisions to local lock files so constraints enforce themselves across sessions and projects.
A Claude Code skill that grounds AI-generated UI in real reference images and a persistent per-developer taste profile, instead of generic AI-slop defaults.
What's new: OpenDots can give each agent its own isolated browser and file workspace that persists across restarts, then wires that computer into a document workspace where agents can draft, review, and save pages with a human-in-the-loop approval step.
Your always-on AI coworkers that move between text, calls, and Slack.
What's new: This repo points to a preprint that gives the formal calculus behind Cordis, grounding dynamic component composition in effect and coeffect theory with a metatheory that shows components can interleave without disturbing one another.
A Programming Paradigm for Spatiotemporal Composability
What's new: VoiceStudio bundles voice cloning, video dubbing, dictation, transcription, and audiobook generation into a single local desktop app with a GUI, a local API, and MCP support, defaulting to the k2-fsa/OmniVoice engine with other engines selectable from a built-in catalog.
VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.
What's new: Where the Dogwood reference interpreter processes an entire event trace at once, this engine ingests events one at a time, checkpoints state to disk, and resumes correctly after a crash.
Implementation of an authorization engine for Dogwood policies
What's new: Dogwood extends Cedar's permit/forbid syntax with temporal predicates (formerly, since, windowed aggregations) so policies can reason over an agent's recent event history, not just the current request.
Reference parser and interpreter for the Dogwood policy language
What's new: CLMs treat the context as a file and give the model unrestricted write access to it, then improve that behavior with evolved natural-language instructions or online RL.
Official repository for "Context Language Models"
What's new: vllmops keeps model configs and project settings in Git-reviewable YAML files and pairs a scriptable CLI with a live Textual TUI, so you can drive the full vLLM lifecycle from both without Docker or a monitoring stack.
Self-hosted vLLM made simple. Git-friendly YAML for models and profiles, full lifecycle from an intuitive CLI and TUI.
What's new: It holds model, tasks, and runtime constant across nine harnesses to isolate harness-level effects on pass rate, cost, cache behavior, and speed, then ships a reproducible skill so a third party can evaluate their own harness on the same task set.
Public results and task definitions for FrontierHarness Eval
What's new: JevBench scores decision models on four axes: Intelligence (chance-corrected accuracy with a penalty below 50), Calibration (ECE and label-distribution fidelity), Speed (log-scale latency at p50 and p95), and Cost reported per 1,000 decisions rather than per 1,000 tokens.
JevBench by Benchmark Heaven: benchmark and leaderboard for AI decision models, measuring typed decision accuracy, calibration, latency and cost.
What's new: It adapts principles inspired by the ISO 24495 plain language series (unofficial, with no conformance claim) into agent-readable skill files and a rule engine that audits Markdown against measurable criteria: sentence length, paragraph length, legalese, heading depth, undefined acronyms, and more.
ISO 24495 Plain Language skills and plugin
What's new: Lighthouse splits its run into a separate gather phase and audit phase, letting you capture browser artifacts once and re-audit them offline without hitting the network again.
Automated auditing, performance metrics, and best practices for the web.
What's new: at_mcp gives the agent ownership of credentials, session, and a write quota, so it holds a persistent AT Protocol identity across runs.
An agent's own AT Protocol account, through MCP
What's new: It ships three interfaces in one repo: a standalone binary for humans, a Streamable HTTP MCP server for agents, and a self-generating Skill file that tells coding agents how to use the CLI commands.
A Command-Line Interface (CLI), Skill and MCP Server to interact with Papers with Code. For agents and humans.
What's new: Unicel stores each cell value as a (number, unit) tuple, making unit cancellation and metric/imperial conversion part of the formula engine rather than a manual step.
Unit-aware spreadsheet application with dimensional analysis
What's new: Alongside the human-readable entries, it maintains a machine-readable policy registry (policies.json) and a revocation list (revocations.json), both versioned in Git so tooling can consume them directly.
A curated list of software, simulators, policies, agent tools and coverage for the Pollen Robotics / Hugging Face Microduck robot
What's new: Kite wraps Jetstream V2 as an OTP-native behaviour module, letting you define a subscriber with compile-time collection filters and pluggable cursor callbacks rather than managing websocket state yourself.
An Elixir atproto Jetstream V2 subscriber library
What's new: The code includes OpenAI-powered scripts for generating idea chains, neologisms, and evaluating ideas, alongside curated text lists on topics like flowers, fractals, and cooperation.
What's new: SwarmWorld measures collective technological evolution at the society level: artifacts persist between ticks, executable programs are inherited by successors, the engine keeps causal knowledge records, and traces are deterministically replayable with counterfactual contributor-removal.
What's new: Trajectory parses raw session transcripts from 15 agent runtimes, each with its own incompatible format, into a single schema-validated record array that training, evaluation, analysis, and inference systems can consume.
Convert sessions across harnesses to a unified trajectory format - designed to be consumed by agents (e.g. for memory formation, dreaming, search)
What's new: pbf splits decoding and encoding into separate tree-shakeable classes and keeps the bundle at 2.5KB gzipped, with lazy decoding and a CLI that compiles `.proto` schemas into plain, editable JavaScript.
A low-level, lightweight protocol buffers implementation in JavaScript.
What's new: Loom derives every REST route, MCP tool, CLI command, and authority boundary from one Rust declaration per operation, so the server, the CLI, and the MCP adapter cannot disagree.
Issue tracking for AI coding agents
What's new: Each skill is a small, composable unit covering one job, so you pick the ones you want and leave the rest.
Skills for Real Engineers. Straight from my .agents directory.
What's new: Rather than wrapping a single analysis engine, REA dispatches across native decompilers, JavaScript/Electron inspection, .NET assembly reading, Android APK tools, firmware extractors, and browser capture through one MCP interface with evidence attached to every finding.
Reverse engineer anything with agents, from app behavior down to native binaries.
What's new: It is based on Patrick Gunkel's Ideonomy and runs as an engine with 28 named lenses, profile-based auto-selection, lens chaining, cross-lens synthesis, and persistent sessions.
Structured creative reasoning for AI agents. Systematic thinking through 28 ideonomic lenses. 🔬
What's new: Superpowers bundles an end-to-end development workflow, from requirements brainstorming through subagent-driven task execution and TDD enforcement, with skills that fire automatically in sequence.
An agentic skills framework & software development methodology that works.
What's new: This pack encodes UX and visual design judgment as agent-readable skills: when to use a specific research method, how to tokenize a design system, and what a heuristic evaluation should cover.
Designer Skills Collection: agentic skills, commands, and plugins for design — from research to systems, UI, interaction, and delivery.
What's new: p5.js carries Processing's setup/draw sketch model directly into the browser, treating the whole web page as your sketch.
p5.js is a client-side JS platform that empowers artists, designers, students, and anyone to learn to code and express themselves creatively on the web. It is based on the core principles of Processing. Looking for p5.js 2.0? http://beta.p5js.org
What's new: It ships a single compiled neural program identified by a fixed ID, downloaded once and then run locally for all subsequent inference.
Detect and type PII locally with one ProgramAsWeights neural program.
What's new: SkyPilot spans 20+ clouds, Kubernetes, and Slurm clusters under a single control plane, routing jobs to available capacity with automatic failover and bin-packing.
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.
What's new: This pack encodes a business methodology, translating Sahil Lavingia's Minimalist Entrepreneur framework into ten sequenced Claude Code commands.
Based on The Minimalist Entrepreneur by Sahil Lavingia
What's new: Box ties OS-level sandboxing to a stateful Dogwood policy engine so a file read through one interpreter can block a later network request, and injects credentials into permitted outbound calls without exposing secrets to the agent process.
Run AI agents in a sandbox that restricts what they can execute, read, write, and reach on the network. Box combines OS isolation with default-deny Dogwood policies and credential injection that keeps secrets outside the agent. Written in Rust. Supports macOS on Apple silicon, with Linux support planned.
What's new: Strands Decider replaces its language-modelling head with a pointer head that scores options by comparing hidden states, producing calibrated confidence values without any generation loop.
A small, fast decision model, or system one model, for agentic workflows. Pick between options or rate on a scale faster than an LLM, with a calibrated confidence on every decision.
What's new: Superset wraps any CLI coding agent in a full desktop IDE with per-worktree browser previews, a diff viewer, scheduled automations, and an MCP server that lets agents create and manage their own workspaces.
Superset is an agentic IDE to orchestrate 100+ coding agents in parallel. Run any agent with your own subscription.
What's new: TiKV pairs classical raw key-value access with full ACID transactional semantics modeled after Google Percolator, giving applications a single store that handles both access patterns.
Distributed transactional key-value database, originally created to complement TiDB
What's new: Rather than a storage backend or vector database, OptMem stores memories as fixed-width append-only lines in a plain text log and builds a binary tree of summaries on top, so the agent navigates and compresses its own history with simple shell commands and no external dependencies.
Permanent memory for AI agents. A 426-token prompt, a script, plug and play.
What's new: OpenViking structures all context as a navigable filesystem under `viking://` URIs with scoped semantic search and three-tier summaries (abstract, overview, full content) so agents can judge relevance before reading.
Self-evolving Context Database for AI Agents. Unify Agent Memory, Knowledge RAG and Skills.
What's new: It tracks a single model's ecosystem by requiring each entry to use Jev, or a documented Jev port or derivative, for a concrete decision task.
A curated list of public projects, integrations, and discussions built on Jev — TypeSafe AI's System One model for typed decisions.