In one paragraph
Data systems cover a wide band: how bytes are laid on the wire, how records are stored and replicated, how data flows between services, and how realistic fake data gets made. pbf handles binary encoding in the browser. tikv stores transactional key-value data across a Raft cluster. airbyte moves data in bulk with 600+ connectors. kite subscribes to a live AT Protocol event stream. sample-synthetic-document-generator turns real PDFs into synthetic copies. What counts here: databases, pipelines, encodings, and generators that store, move, or produce data.
The main approaches
Encodings
Pack data into a compact binary format for fast transfer or storage. The repo compiles a schema into typed read/write functions and handles the wire format so application code does not have to. Example: mapbox/pbf.
Databases
Store and retrieve records with transactional guarantees, replication, and horizontal scaling. The system manages sharding, consensus, and failure recovery. Example: tikv/tikv.
Data movement
Connect a source to a destination, transforming records along the way. This covers both bulk batch loads (ELT pipelines) and live event streams. Examples: airbytehq/airbyte, joladev/kite.
Synthetic data
Generate realistic but fictional data from a real document or schema so teams can test, demo, or share data without exposing production records. Example: aws-samples/sample-synthetic-document-generator.
Map of the theme
flowchart LR t["Data systems"] t --> f1["Databases"] t --> f2["Data movement"] t --> f3["Encodings"] t --> f4["Synthetic data"] f1 --> r1["tikv/tikv"] f2 --> r2["airbytehq/airbyte"] f2 --> r3["joladev/kite"] f3 --> r4["mapbox/pbf"] f4 --> r5["aws-samples/sample-synthetic-document-generator"]
Where the new ideas are
- mapbox/pbf: tree-shakeable split between decoder and encoder classes, keeping the bundle at 2.5KB gzipped while a CLI compiles
.protoschemas to plain, editable JavaScript. - tikv/tikv: layering full Percolator-style ACID transactions on top of a raw key-value store so both access patterns share one backend.
- airbytehq/airbyte: a separate Agent SDK that exposes connectors as typed LLM tools for pydantic-ai, LangChain, OpenAI Agents, and FastMCP.
- joladev/kite: OTP-native behaviour module for AT Protocol Jetstream V2 with pluggable cursor persistence, so restart recovery is a callback rather than manual websocket state management.
- aws-samples/sample-synthetic-document-generator: separating paid schema extraction (Bedrock Claude) from free, offline, seeded row generation, and adding an Amazon Comprehend PII audit pass over the output.
Side by side
| Repo | Approach | Language | Schema/Format support | Distributed / multi-node | Offline / self-contained operation |
|---|---|---|---|---|---|
| mapbox/pbf | Encodings | JavaScript | .proto compiled to JS via CLI | Not stated | Yes (browser and Node) |
| tikv/tikv | Databases | Rust | Key-value; raw and transactional APIs | Yes, Raft + Placement Driver, 100+ TB | Yes (self-hosted) |
| airbytehq/airbyte | Data movement | Python | 600+ source/destination connectors | Not stated | Yes (self-hosted) |
| aws-samples/sample-synthetic-document-generator | Synthetic data | Python | PDF in, HTML/CSV/JSON out; 15 bundled presets | Not stated | Partial (schema extraction needs Bedrock; row generation is offline) |
| joladev/kite | Data movement | Elixir | AT Protocol Jetstream V2 events (NSID filter) | Client-side failover across endpoint list | Not stated |
How the idea moved
flowchart LR n1["started Jan 2014<br/>mapbox/pbf"] n2["started Dec 2015<br/>tikv/tikv"] n3["started Jul 2020<br/>airbytehq/airbyte"] n4["started May 2026<br/>aws-samples/sample-synthetic-document-generator"] n5["started Sep 2026<br/>joladev/kite"] n1 --> n2 --> n3 --> n4 --> n5
- started Jan 2014 · pbf · Adds: Provides a 2.5KB gzipped protobuf codec for JavaScript with a CLI that compiles
.protoschemas into plain, tree-shakeable read/write modules, benchmarked at 192 MB/s decode and 257 MB/s encode for Mapbox vector tile data. - started Dec 2015 · tikv · Adds: Delivers a Raft-replicated, RocksDB-backed key-value store with Percolator-style ACID transactions, snapshot isolation, and a Placement Driver that handles sharding and rebalancing up to 100+ TB.
- started Jul 2020 · airbyte · Adds: Connects 600+ sources and destinations through a no-code/low-code connector framework and adds a separate Agent SDK that turns connectors into LLM tools for pydantic-ai, LangChain, OpenAI Agents, and FastMCP.
- started May 2026 · sample-synthetic-document-generator · Adds: Rewrites a reference PDF into synthetic HTML copies with Claude on Bedrock, runs an Amazon Comprehend PII audit on the output, and separates paid Bedrock schema extraction from free offline seeded CSV/JSON row generation.
- started Sep 2026 · kite · Adds: Wraps AT Protocol Jetstream V2 as an OTP behaviour with compile-time collection filters, pluggable cursor persistence for restart recovery, and automatic failover across a configurable endpoint list.
Easily confused
This theme might be confused with general developer tooling or API clients. The difference is scope: everything here is specifically about storing, moving, or producing data as its primary job. A web framework that happens to query a database is not a data system. An ORM is not a data system. The stores, the wires, the codecs, and the generators belong here; the application layer that sits above them does not.
Gaps nobody has filled
- End-to-end data latency measurement from source event to destination write is absent across all approaches.
- Schema evolution is handled tool by tool: Airbyte detects and propagates schema changes per connection and pbf compiles each
.protoon its own, and no member offers a schema registry shared across them. - Synthetic data generation covers documents and tabular rows but does not produce realistic graph or time-series data.
- Kite reads only AT Protocol Jetstream. Airbyte runs scheduled syncs, even for CDC sources, and no member is an OTP-native subscriber for event buses such as Kafka or Kinesis.
- Data lineage or audit trails that track a record from its origin through transformations to its final destination are not addressed.