Vitalik Buterin Tests a Privacy-First AI Stack: Local Qwen, zkAPI, and Tor
On Oct 4 Vitalik tested privacy AI with local Qwen3.8-Flash-Next, zkAPI payments, and Tor routing—20-30 tokens/sec locally, Tor adds heavy latency.
7 min read
On October 4, 2026, Vitalik Buterin published results from a personal experiment that reads like a systems integration test for the next decade of AI privacy. Using a local Alibaba Qwen3.8-Flash-Next model, Ethereum's newly launched zkAPI payment layer, and Tor network routing, he generated personalized diet and exercise recommendations from sensitive health and travel data while attempting to prevent remote services from linking his identity, his wallet, or his network footprint. The setup worked end to end. It was also slow, constrained, and honest about tradeoffs—exactly the kind of prototype serious privacy engineering produces before UX polish arrives.
Buterin disclosed the test on Farcaster, framing it not as a product launch but as a research exercise in composability. Each layer solves a different leakage class. Local inference handles identity-shaped metadata embedded in how humans write prompts. zkAPI handles payment-shaped metadata that would otherwise tie API keys to accounts. Tor handles network-shaped metadata like IP addresses and coarse timing observables at the transport layer. None of the three alone delivers "private AI" in any absolute sense; together they sketch a reference architecture developers can iterate on as models get faster and proving systems get cheaper.
Layer one: local Qwen as a prompt firewall
The local model in Buterin's stack is Qwen3.8-Flash-Next, a compact variant in Alibaba's Qwen family tuned for efficient on-device or on-server inference. Its job is not primarily to produce final health advice—though it may contribute to that—but to decide what information a remote frontier model truly needs and to rewrite requests before they leave the trust boundary.
That architectural choice reflects a lesson privacy engineers learned long ago in federated and split-inference systems: the biggest leaks are often stylistic and contextual, not literal name fields. A user who mentions a specific hotel gym, a medication schedule, and a distinctive rhetorical habit gives a remote provider enough signal to re-identify or persistently profile even if the payment rail is anonymous. By interposing a local rewriter, Buterin aims to strip or generalize identifying flourishes while preserving task utility.
Performance numbers from his test ground the ambition in 2026 hardware reality. The local model ran at roughly twenty to thirty tokens per second. Buterin noted that a fluid conversational experience would want something closer to a hundred tokens per second or higher—implying either better quantization, stronger local GPUs, or smaller task-specific adapters. Privacy here costs latency and capital equipment, not just subscription fees.
Layer two: zkAPI separates payments from prompts
The payment layer uses zkAPI, which the Ethereum Foundation and Open Anonymity Project placed on mainnet on October 1, 2026. Buterin's experiment sits among the first high-profile uses cases validating the protocol outside demo chat iframes.
In zkAPI, users deposit ETH or USDC into an on-chain vault and hold a private note locally. Spending authorizations are Groth16 zero-knowledge proofs that demonstrate funded, unspent state without revealing which note paid for which request. For OpenRouter-compatible flows, a prompt-free lease pattern lets the browser obtain short-lived API keys and speak directly to model endpoints while the billing server settles measured usage afterward.
Buterin's health workflow benefits from that separation because remote model providers often log account identifiers, payment tokens, and request payloads in one observability pane. zkAPI breaks the financial thread. It does not erase prompt content from provider logs—a point the Foundation's own documentation stresses—and it does not hide that a given IP address at a given time spoke to an API. Hence layer three.
Layer three: Tor and the latency tax
Tor routes traffic through multiple relays to obscure origin IP addresses from destination services. Buterin identified Tor as the weakest link in his current stack—not because it fails cryptographically, but because it was not engineered for per-request unlinkability at AI interaction frequencies.
His testing reported latency multipliers roughly between ten and one hundred times compared to direct connections. For interactive chat, that is often unacceptable; for batch health planning where answers can arrive minutes later, it may be tolerable. GitHub activity around zkAPI includes an open pull request—numbered 16 in public references at the time of writing—to add Tor-routed client support, adjusting timeouts to accommodate slower circuits. That kind of patch is mundane and essential: privacy stacks fail in production when defaults assume data-center RTTs.
Buterin also argued that all three protections are necessary because threat models compose. A provider who cannot see your wallet might still recognize your writing; a network adversary who cannot see your IP might still correlate payment deposits on-chain; a local model alone cannot access frontier capabilities without remote calls. Defense in depth is not paranoia—it is acknowledgment that AI privacy is multi-dimensional.
The health use case as a stress test
Diet and exercise recommendations sound benign compared to political or medical diagnostic chat, but they are information-dense. They involve biometrics, schedules, location hints, cultural food preferences, and travel constraints. They are also exactly the sort of queries users might hesitate to type into a logged SaaS account tied to their real email.
Buterin's experiment asks whether personalized outcomes require personalized identifiers. The local rewriter hypothesis says no—or at least, not as many identifiers as users routinely leak. The remote frontier model still adds value for reasoning and breadth, but only after minimization. The stricter the minimization, Buterin noted, the less helpful the remote model can be—a fundamental tradeoff curve privacy product managers will recognize from differential privacy and data minimization regimes in regulated industries.
He also contributed code toward the zkAPI repository, signaling that this was hands-on engineering, not a theoretical thread. Open-source iteration matters because privacy tooling that cannot be reproduced and audited independently rarely earns trust from security-sensitive communities.
Three shortcomings named explicitly
Buterin listed three limitations with engineer's clarity.
First, Tor's efficiency for request-by-request disassociation is poor relative to ideal anonymity systems purpose-built for high-churn API clients. Second, local model throughput at twenty to thirty tokens per second is below smooth-interaction thresholds without hardware upgrades or model distillation. Third, stronger data protection reduces the remote model's actionable context—an economic and utility cost, not a bug.
That honesty is strategically useful. Privacy marketing often overpromises; Buterin's post underpromises and shows working code. For developers, the actionable insight is to benchmark each layer independently before compositing. Measure rewriter fidelity locally. Measure proof generation time for zkAPI authorizations. Measure Tor latency distributions for your target regions. Only then decide which queries belong in a high-privacy pipeline versus a conventional account-based app.
How this fits Ethereum's AI narrative
Buterin has argued for months that Ethereum's role in AI is likelier to be cryptographic plumbing—payments, commitments, identity boundaries—than running trillion-parameter models on-chain. zkAPI embodies that thesis. His personal health stack is a vertical slice: Ethereum secures prepaid credits without account linkage; off-chain models do the inference; Tor optionally hardens network observability.
The experiment also intersects with Glamsterdam-era infrastructure conversations happening the same week on testnets, where throughput and gas pricing upgrades prepare Ethereum for more complex execution patterns. Privacy-preserving AI will not consume chain bandwidth for inference, but it will consume it for deposits, withdrawals, and proof verification unless L2 abstractions mature.
What comes next for builders
Teams inspired by Buterin's stack should treat it as a pattern, not a product. Ingredients include a capable small local model, a remote API with clear metering, a ZK prepaid rail, and optional Tor with session-isolated circuits. Missing pieces for mainstream adoption include hardware profiles that hit hundred-plus tokens per second locally, multiparty trusted setups for proving keys, standardized prompt minimization benchmarks, and UX that hides nullifier state management from non-cryptographers.
Pull request sixteen's Tor client work is a concrete next step in the zkAPI repo. Wider community contributions might explore VPN alternatives for lower latency, confidential compute enclaves for remote inference without raw prompt exposure, and policy languages that let users cap spend and sensitivity per task.
Buterin did not claim victory over surveillance capitalism in a Farcaster post. He claimed a working prototype with known flaws—arguably more valuable. In October 2026, privacy-first AI is not a single switch. It is local Qwen rewriting your words, zkAPI unlinking your wallet, and Tor slowing your packets so a remote model can help without knowing who paid or where you slept last night. Imperfect, measurable, and worth building on.
More in cryptocurrency
Cubed
Write about the technologies shaping the future.
For developers, founders, and curious minds exploring AI, crypto, Web3, and emerging tech—signal over noise.
One free account across In Plain English, Stackademic, Venture, and Cubed.
How it works- AI, crypto & Web3
- Software & emerging technologies
- Analysis & practical resources
- Thoughtful voices, not hype
Sign in
Google or GitHub
Complete profile
Takes a few minutes
Get approved & publish
Start sharing
Why write for Cubed?
The future deserves thoughtful voices, not just louder headlines.

Comments
Loading comments…