NDJSON Explained for Streaming APIs and LLM Pipelines
Newline-delimited JSON (NDJSON) is everywhere in AI and data tooling. Here is what it is, why teams use it, and how to parse it safely.
6 min read
If you have worked with streaming LLM APIs, log shippers, or bulk data exports lately, you have almost certainly touched NDJSON whether you knew the name or not. NDJSON—newline-delimited JSON, also called JSON Lines or .jsonl—is a simple format with outsized impact on how modern systems move data.
Each line in an NDJSON file is one complete JSON value, usually an object. Lines are separated by \n. There is no wrapping array, no trailing commas between records, and no requirement that every object share the same keys.
{"id": 1, "event": "login", "user": "alice"}
{"id": 2, "event": "purchase", "amount": 49.99}
{"id": 3, "event": "logout", "user": "alice"}
That is the entire format. Its power comes from what this structure enables.
Why NDJSON exists
Classic JSON is great for single documents. It is awkward for unbounded streams because a valid JSON file must be one value—typically an array of objects wrapped in [ and ].
That array wrapper creates problems at scale:
- You cannot append to a JSON array file without parsing and rewriting the whole thing
- Parsers must load the entire array into memory
- A corrupted byte in the middle can invalidate the entire file
- Streaming producers and consumers cannot work line by line
NDJSON fixes this by making each record self-contained. Producers append a line. Consumers read line by line. A bad line might fail locally without destroying the rest of the file.
Where you will see NDJSON in the wild
LLM fine-tuning datasets. OpenAI, Hugging Face, and many training pipelines expect .jsonl files where each line is a training example—often with messages arrays for chat formats or prompt/completion pairs for legacy setups.
Streaming API responses. Some endpoints emit partial results as NDJSON so clients render tokens or progress events incrementally instead of waiting for a complete JSON blob.
Log aggregation. Tools like Elasticsearch and cloud logging pipelines ingest JSON-per-line logs. NDJSON is a natural fit for append-only log files.
ETL and data lakes. Export jobs from databases and SaaS products often land in S3 as .jsonl because Spark, BigQuery, and AWS Glue handle line-delimited files efficiently.
Bulk API operations. When you need to POST thousands of create/update operations, vendors sometimes accept an NDJSON upload rather than a giant array payload.
If you are building AI tooling, knowing NDJSON is as practical as knowing HTTP status codes.
NDJSON vs JSON vs CSV
| Format | Strength | Weakness |
|---|---|---|
| NDJSON | Streamable, append-friendly, schema-flexible per row | No single-schema enforcement; redundant keys |
| JSON array | One parse gives full dataset | Poor for huge or live streams |
| CSV | Compact for flat tabular data | Nested structures become painful |
NDJSON shines when records are semi-structured and arrive continuously. CSV still wins for simple spreadsheets. JSON arrays still win for small config files you edit by hand.
Parsing NDJSON safely
JavaScript (Node)
import readline from "node:readline";
import fs from "node:fs";
const rl = readline.createInterface({
input: fs.createReadStream("data.jsonl"),
crlfDelay: Infinity,
});
for await (const line of rl) {
if (!line.trim()) continue;
const record = JSON.parse(line);
// process record
}
Never JSON.parse the entire file as one value unless you know it is tiny.
Python
import json
with open("data.jsonl", encoding="utf-8") as f:
for line_num, line in enumerate(f, start=1):
line = line.strip()
if not line:
continue
try:
record = json.loads(line)
except json.JSONDecodeError as exc:
raise ValueError(f"Invalid JSON on line {line_num}") from exc
# process record
For large files, this loop keeps memory flat.
Validating streams
Production pipelines should:
- Skip or quarantine empty lines
- Catch
JSONDecodeErrorper line and route bad lines to a dead-letter file - Enforce size limits per line to avoid memory bombs
- Validate schema per record with JSON Schema, Pydantic, or Zod—even if keys vary row to row
Writing NDJSON
Append one serialized object per line:
import json
def write_record(handle, obj):
handle.write(json.dumps(obj, ensure_ascii=False) + "\n")
Use ensure_ascii=False when you expect unicode text. Always include the trailing newline.
When generating training files for LLMs, keep UTF-8 encoding consistent and avoid BOM headers unless your downstream tool expects them.
NDJSON in LLM pipelines
Fine-tuning example shape
Chat-style fine-tuning often looks like:
{"messages": [{"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "Summarize NDJSON."}, {"role": "assistant", "content": "NDJSON is..."}]}
Each line is one conversation. Mixing multiple conversations on one line is invalid.
Evaluation outputs
Batch inference scripts frequently write model outputs as NDJSON so you can resume after failures:
{"id": "doc-991", "model": "gpt-4.1", "output": "...", "latency_ms": 842}
Downstream analysis tools filter, join, and aggregate without reloading everything into RAM.
RAG ingestion
Some chunking pipelines emit one JSON object per chunk with metadata:
{"doc_id": "handbook-v3", "chunk_index": 12, "text": "...", "embedding": [0.01, -0.02, ...]}
Vector databases and search indexes consume these files in bulk import jobs.
Streaming HTTP and NDJSON
Servers can set Content-Type: application/x-ndjson or application/jsonlines and flush one JSON object per line as events occur. Clients read the body as a stream.
This pattern appears in:
- Server-sent style LLM token streams (alongside SSE in other APIs)
- Progress reporting for long-running batch jobs
- Incremental search results
Clients must handle partial lines at TCP chunk boundaries. Buffer incomplete lines until a \n arrives before calling JSON.parse.
Common mistakes
Pretty-printing across lines. Multi-line JSON objects break NDJSON parsers. Pretty-print each record internally if needed, but one logical record must occupy one physical line—or escape newlines inside string values only.
Trailing commas. JSON does not allow trailing commas after the last property. Invalid lines fail ingestion.
Mixing encodings. Stick to UTF-8 end to end.
Assuming homogeneous schemas. Line 1 might have "label" while line 2 does not. Downstream code should use optional fields or schema validation.
Giant single-line payloads. A "line" that is 200 MB defeats streaming benefits. Chunk large binary data by reference (S3 URL) instead of embedding base64 in JSONL.
Tools that speak NDJSON
jq— filter and transform withjq -cto compact each output line- Unix
split— partition large files by line count - Apache Spark —
spark.read.json("file.jsonl")treats each line as a row - BigQuery — load JSONL from GCS with autodetected schema
- Hugging Face
datasets— load JSON lines for training experiments
For quick inspection: head -n 3 file.jsonl | jq .
Security notes
NDJSON files often contain user-generated content. Treat them like any untrusted input:
- Do not eval or execute string fields
- Sanitize before rendering in HTML
- Scan uploads for zip bombs disguised as text
- Restrict file permissions on shared buckets—
.jsonlexports may include PII
When exposing NDJSON downloads publicly, remember search engines and crawlers may index URLs if they are not authenticated.
FAQ
Is NDJSON standardized?
There is no single RFC, but the JSON Lines spec and widespread tool support make it a de facto standard. Names vary—JSONL, LDJSON, NDJSON—but the newline-separated rule is consistent.
Can I gzip NDJSON?
Yes. .jsonl.gz is common for storage and transfer. Decompress before line parsing, or use streaming gzip libraries.
How do I convert CSV to NDJSON?
Read CSV rows, map each to a dict, json.dumps per row. Nested fields require conventions (JSON strings in cells or separate normalization step).
Why not Parquet instead?
Parquet is columnar and efficient for analytics warehouses. NDJSON is simpler for human debugging, LLM tooling, and line-at-a-time streaming. Many pipelines use both—NDJSON for ingest debuggability, Parquet for warehouse storage.
NDJSON will not win design awards. It is boring on purpose—and boring formats scale. If you are wiring together APIs, training data, or log pipelines in 2026, getting comfortable with newline-delimited JSON pays off quickly.
More in apis
Cubed
Write about the technologies shaping the future.
For developers, founders, and curious minds exploring AI, crypto, Web3, and emerging tech—signal over noise.
One free account across In Plain English, Stackademic, Venture, and Cubed.
How it works- AI, crypto & Web3
- Software & emerging technologies
- Analysis & practical resources
- Thoughtful voices, not hype
Sign in
Google or GitHub
Complete profile
Takes a few minutes
Get approved & publish
Start sharing
Why write for Cubed?
The future deserves thoughtful voices, not just louder headlines.
Comments
Loading comments…