NDJSON Explained for Streaming APIs and LLM Pipelines

Newline-delimited JSON (NDJSON) is everywhere in AI and data tooling. Here is what it is, why teams use it, and how to parse it safely.

6 min read

If you have worked with streaming LLM APIs, log shippers, or bulk data exports lately, you have almost certainly touched NDJSON whether you knew the name or not. NDJSON—newline-delimited JSON, also called JSON Lines or .jsonl—is a simple format with outsized impact on how modern systems move data.

Each line in an NDJSON file is one complete JSON value, usually an object. Lines are separated by \n. There is no wrapping array, no trailing commas between records, and no requirement that every object share the same keys.

{"id": 1, "event": "login", "user": "alice"}
{"id": 2, "event": "purchase", "amount": 49.99}
{"id": 3, "event": "logout", "user": "alice"}

That is the entire format. Its power comes from what this structure enables.

Why NDJSON exists

Classic JSON is great for single documents. It is awkward for unbounded streams because a valid JSON file must be one value—typically an array of objects wrapped in [ and ].

That array wrapper creates problems at scale:

  • You cannot append to a JSON array file without parsing and rewriting the whole thing
  • Parsers must load the entire array into memory
  • A corrupted byte in the middle can invalidate the entire file
  • Streaming producers and consumers cannot work line by line

NDJSON fixes this by making each record self-contained. Producers append a line. Consumers read line by line. A bad line might fail locally without destroying the rest of the file.

Where you will see NDJSON in the wild

LLM fine-tuning datasets. OpenAI, Hugging Face, and many training pipelines expect .jsonl files where each line is a training example—often with messages arrays for chat formats or prompt/completion pairs for legacy setups.

Streaming API responses. Some endpoints emit partial results as NDJSON so clients render tokens or progress events incrementally instead of waiting for a complete JSON blob.

Log aggregation. Tools like Elasticsearch and cloud logging pipelines ingest JSON-per-line logs. NDJSON is a natural fit for append-only log files.

ETL and data lakes. Export jobs from databases and SaaS products often land in S3 as .jsonl because Spark, BigQuery, and AWS Glue handle line-delimited files efficiently.

Bulk API operations. When you need to POST thousands of create/update operations, vendors sometimes accept an NDJSON upload rather than a giant array payload.

If you are building AI tooling, knowing NDJSON is as practical as knowing HTTP status codes.

NDJSON vs JSON vs CSV

FormatStrengthWeakness
NDJSONStreamable, append-friendly, schema-flexible per rowNo single-schema enforcement; redundant keys
JSON arrayOne parse gives full datasetPoor for huge or live streams
CSVCompact for flat tabular dataNested structures become painful

NDJSON shines when records are semi-structured and arrive continuously. CSV still wins for simple spreadsheets. JSON arrays still win for small config files you edit by hand.

Parsing NDJSON safely

JavaScript (Node)

import readline from "node:readline";
import fs from "node:fs";

const rl = readline.createInterface({
  input: fs.createReadStream("data.jsonl"),
  crlfDelay: Infinity,
});

for await (const line of rl) {
  if (!line.trim()) continue;
  const record = JSON.parse(line);
  // process record
}

Never JSON.parse the entire file as one value unless you know it is tiny.

Python

import json

with open("data.jsonl", encoding="utf-8") as f:
    for line_num, line in enumerate(f, start=1):
        line = line.strip()
        if not line:
            continue
        try:
            record = json.loads(line)
        except json.JSONDecodeError as exc:
            raise ValueError(f"Invalid JSON on line {line_num}") from exc
        # process record

For large files, this loop keeps memory flat.

Validating streams

Production pipelines should:

  • Skip or quarantine empty lines
  • Catch JSONDecodeError per line and route bad lines to a dead-letter file
  • Enforce size limits per line to avoid memory bombs
  • Validate schema per record with JSON Schema, Pydantic, or Zod—even if keys vary row to row

Writing NDJSON

Append one serialized object per line:

import json

def write_record(handle, obj):
    handle.write(json.dumps(obj, ensure_ascii=False) + "\n")

Use ensure_ascii=False when you expect unicode text. Always include the trailing newline.

When generating training files for LLMs, keep UTF-8 encoding consistent and avoid BOM headers unless your downstream tool expects them.

NDJSON in LLM pipelines

Fine-tuning example shape

Chat-style fine-tuning often looks like:

{"messages": [{"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "Summarize NDJSON."}, {"role": "assistant", "content": "NDJSON is..."}]}

Each line is one conversation. Mixing multiple conversations on one line is invalid.

Evaluation outputs

Batch inference scripts frequently write model outputs as NDJSON so you can resume after failures:

{"id": "doc-991", "model": "gpt-4.1", "output": "...", "latency_ms": 842}

Downstream analysis tools filter, join, and aggregate without reloading everything into RAM.

RAG ingestion

Some chunking pipelines emit one JSON object per chunk with metadata:

{"doc_id": "handbook-v3", "chunk_index": 12, "text": "...", "embedding": [0.01, -0.02, ...]}

Vector databases and search indexes consume these files in bulk import jobs.

Streaming HTTP and NDJSON

Servers can set Content-Type: application/x-ndjson or application/jsonlines and flush one JSON object per line as events occur. Clients read the body as a stream.

This pattern appears in:

  • Server-sent style LLM token streams (alongside SSE in other APIs)
  • Progress reporting for long-running batch jobs
  • Incremental search results

Clients must handle partial lines at TCP chunk boundaries. Buffer incomplete lines until a \n arrives before calling JSON.parse.

Common mistakes

Pretty-printing across lines. Multi-line JSON objects break NDJSON parsers. Pretty-print each record internally if needed, but one logical record must occupy one physical line—or escape newlines inside string values only.

Trailing commas. JSON does not allow trailing commas after the last property. Invalid lines fail ingestion.

Mixing encodings. Stick to UTF-8 end to end.

Assuming homogeneous schemas. Line 1 might have "label" while line 2 does not. Downstream code should use optional fields or schema validation.

Giant single-line payloads. A "line" that is 200 MB defeats streaming benefits. Chunk large binary data by reference (S3 URL) instead of embedding base64 in JSONL.

Tools that speak NDJSON

  • jq — filter and transform with jq -c to compact each output line
  • Unix split — partition large files by line count
  • Apache Spark — spark.read.json("file.jsonl") treats each line as a row
  • BigQuery — load JSONL from GCS with autodetected schema
  • Hugging Face datasets — load JSON lines for training experiments

For quick inspection: head -n 3 file.jsonl | jq .

Security notes

NDJSON files often contain user-generated content. Treat them like any untrusted input:

  • Do not eval or execute string fields
  • Sanitize before rendering in HTML
  • Scan uploads for zip bombs disguised as text
  • Restrict file permissions on shared buckets—.jsonl exports may include PII

When exposing NDJSON downloads publicly, remember search engines and crawlers may index URLs if they are not authenticated.

FAQ

Is NDJSON standardized?
There is no single RFC, but the JSON Lines spec and widespread tool support make it a de facto standard. Names vary—JSONL, LDJSON, NDJSON—but the newline-separated rule is consistent.

Can I gzip NDJSON?
Yes. .jsonl.gz is common for storage and transfer. Decompress before line parsing, or use streaming gzip libraries.

How do I convert CSV to NDJSON?
Read CSV rows, map each to a dict, json.dumps per row. Nested fields require conventions (JSON strings in cells or separate normalization step).

Why not Parquet instead?
Parquet is columnar and efficient for analytics warehouses. NDJSON is simpler for human debugging, LLM tooling, and line-at-a-time streaming. Many pipelines use both—NDJSON for ingest debuggability, Parquet for warehouse storage.

NDJSON will not win design awards. It is boring on purpose—and boring formats scale. If you are wiring together APIs, training data, or log pipelines in 2026, getting comfortable with newline-delimited JSON pays off quickly.

More in apis

Cubed

Write about the technologies shaping the future.

For developers, founders, and curious minds exploring AI, crypto, Web3, and emerging tech—signal over noise.

One free account across In Plain English, Stackademic, Venture, and Cubed.

How it works
  • AI, crypto & Web3
  • Software & emerging technologies
  • Analysis & practical resources
  • Thoughtful voices, not hype
1

Sign in

Google or GitHub

2

Complete profile

Takes a few minutes

3

Get approved & publish

Start sharing

Why write for Cubed?

The future deserves thoughtful voices, not just louder headlines.

Comments

Loading comments…

Posts Across the Network