Theme

Data systems

5 repos reviewed · since Oct 10, 2026 · updated Oct 10, 2026 · rewritten each time a repo joins

In one paragraph

Data systems cover a wide band: how bytes are laid on the wire, how records are stored and replicated, how data flows between services, and how realistic fake data gets made. pbf handles binary encoding in the browser. tikv stores transactional key-value data across a Raft cluster. airbyte moves data in bulk with 600+ connectors. kite subscribes to a live AT Protocol event stream. sample-synthetic-document-generator turns real PDFs into synthetic copies. What counts here: databases, pipelines, encodings, and generators that store, move, or produce data.

The main approaches

Encodings

Pack data into a compact binary format for fast transfer or storage. The repo compiles a schema into typed read/write functions and handles the wire format so application code does not have to. Example: mapbox/pbf.

Databases

Store and retrieve records with transactional guarantees, replication, and horizontal scaling. The system manages sharding, consensus, and failure recovery. Example: tikv/tikv.

Data movement

Connect a source to a destination, transforming records along the way. This covers both bulk batch loads (ELT pipelines) and live event streams. Examples: airbytehq/airbyte, joladev/kite.

Synthetic data

Generate realistic but fictional data from a real document or schema so teams can test, demo, or share data without exposing production records. Example: aws-samples/sample-synthetic-document-generator.

Map of the theme

flowchart LR
  t["Data systems"]
  t --> f1["Databases"]
  t --> f2["Data movement"]
  t --> f3["Encodings"]
  t --> f4["Synthetic data"]
  f1 --> r1["tikv/tikv"]
  f2 --> r2["airbytehq/airbyte"]
  f2 --> r3["joladev/kite"]
  f3 --> r4["mapbox/pbf"]
  f4 --> r5["aws-samples/sample-synthetic-document-generator"]

Where the new ideas are

  • mapbox/pbf: tree-shakeable split between decoder and encoder classes, keeping the bundle at 2.5KB gzipped while a CLI compiles .proto schemas to plain, editable JavaScript.
  • tikv/tikv: layering full Percolator-style ACID transactions on top of a raw key-value store so both access patterns share one backend.
  • airbytehq/airbyte: a separate Agent SDK that exposes connectors as typed LLM tools for pydantic-ai, LangChain, OpenAI Agents, and FastMCP.
  • joladev/kite: OTP-native behaviour module for AT Protocol Jetstream V2 with pluggable cursor persistence, so restart recovery is a callback rather than manual websocket state management.
  • aws-samples/sample-synthetic-document-generator: separating paid schema extraction (Bedrock Claude) from free, offline, seeded row generation, and adding an Amazon Comprehend PII audit pass over the output.

Side by side

RepoApproachLanguageSchema/Format supportDistributed / multi-nodeOffline / self-contained operation
mapbox/pbfEncodingsJavaScript.proto compiled to JS via CLINot statedYes (browser and Node)
tikv/tikvDatabasesRustKey-value; raw and transactional APIsYes, Raft + Placement Driver, 100+ TBYes (self-hosted)
airbytehq/airbyteData movementPython600+ source/destination connectorsNot statedYes (self-hosted)
aws-samples/sample-synthetic-document-generatorSynthetic dataPythonPDF in, HTML/CSV/JSON out; 15 bundled presetsNot statedPartial (schema extraction needs Bedrock; row generation is offline)
joladev/kiteData movementElixirAT Protocol Jetstream V2 events (NSID filter)Client-side failover across endpoint listNot stated

How the idea moved

flowchart LR
  n1["started Jan 2014<br/>mapbox/pbf"]
  n2["started Dec 2015<br/>tikv/tikv"]
  n3["started Jul 2020<br/>airbytehq/airbyte"]
  n4["started May 2026<br/>aws-samples/sample-synthetic-document-generator"]
  n5["started Sep 2026<br/>joladev/kite"]
  n1 --> n2 --> n3 --> n4 --> n5
  • started Jan 2014 · pbf · Adds: Provides a 2.5KB gzipped protobuf codec for JavaScript with a CLI that compiles .proto schemas into plain, tree-shakeable read/write modules, benchmarked at 192 MB/s decode and 257 MB/s encode for Mapbox vector tile data.
  • started Dec 2015 · tikv · Adds: Delivers a Raft-replicated, RocksDB-backed key-value store with Percolator-style ACID transactions, snapshot isolation, and a Placement Driver that handles sharding and rebalancing up to 100+ TB.
  • started Jul 2020 · airbyte · Adds: Connects 600+ sources and destinations through a no-code/low-code connector framework and adds a separate Agent SDK that turns connectors into LLM tools for pydantic-ai, LangChain, OpenAI Agents, and FastMCP.
  • started May 2026 · sample-synthetic-document-generator · Adds: Rewrites a reference PDF into synthetic HTML copies with Claude on Bedrock, runs an Amazon Comprehend PII audit on the output, and separates paid Bedrock schema extraction from free offline seeded CSV/JSON row generation.
  • started Sep 2026 · kite · Adds: Wraps AT Protocol Jetstream V2 as an OTP behaviour with compile-time collection filters, pluggable cursor persistence for restart recovery, and automatic failover across a configurable endpoint list.

Easily confused

This theme might be confused with general developer tooling or API clients. The difference is scope: everything here is specifically about storing, moving, or producing data as its primary job. A web framework that happens to query a database is not a data system. An ORM is not a data system. The stores, the wires, the codecs, and the generators belong here; the application layer that sits above them does not.

Gaps nobody has filled

  • End-to-end data latency measurement from source event to destination write is absent across all approaches.
  • Schema evolution is handled tool by tool: Airbyte detects and propagates schema changes per connection and pbf compiles each .proto on its own, and no member offers a schema registry shared across them.
  • Synthetic data generation covers documents and tabular rows but does not produce realistic graph or time-series data.
  • Kite reads only AT Protocol Jetstream. Airbyte runs scheduled syncs, even for CDC sources, and no member is an OTP-native subscriber for event buses such as Kafka or Kinesis.
  • Data lineage or audit trails that track a record from its origin through transformations to its final destination are not addressed.