Skip to content
Back to Insights

Optimising Neo4J Bulk Import

Lessons from loading billion-node graphs – trading off speed, cost, and data quality.

Saurabh Mehrotra
Saurabh Mehrotra
Director at Xpergia

If you’ve worked with Neo4J’s bulk import tool on anything beyond a toy dataset, you’ll know that the defaults don’t cut it. I’ve spent a fair amount of time tuning imports for large graphs – billions of nodes and relationships – and along the way I’ve picked up a few lessons about the trade-offs that I wish someone had written down for me earlier.

This post shares my perspective on getting the most out of the bulk import tool. For the official reference on available options and syntax, I’d recommend starting with the Neo4j bulk import documentation.

A note on versions: If you’re on Neo4j 5.x, the CLI syntax has changed. The import tool now lives under neo4j-admin database import full and neo4j-admin database import incremental. The options I discuss below apply to this newer structure. Always check the official docs for the exact syntax for your version.

Why Full Loads Are Still the Norm

The incremental import option introduced in Neo4j 5 is a welcome addition, but in my experience it hasn’t replaced full loads for most use cases. The reason is simple: incremental mode only lets you append new nodes and relationships. You can’t update existing properties or delete anything.

Choosing Your Import Strategy Need to update or delete data? YES FULL LOAD Rebuilds entire database NO INCREMENTAL Appends only – no updates ⚠ DB must be stopped ⚠ Use --overwrite-destination ✓ Updates & deletes ✓ Full restructuring ✗ Expensive at scale ✓ DB can stay running ✓ Fast for additions ✓ Minimal downtime ✗ No updates/deletes ✗ No restructuring
Decision flow: full load covers all cases but is expensive; incremental is fast but append-only.

In the real world, graph data evolves – nodes change, relationships are restructured, properties get corrected. That almost always means a full load.

Two things caught me off guard early on. First, if you’re running iterative full loads (which you almost certainly will during development), make sure you use the --overwrite-destination flag – without it, the import refuses to write to an existing database. Second, the target database needs to be stopped during import – easy to overlook when planning maintenance windows.

neo4j-admin database import full \
  --overwrite-destination \
  --nodes=nodes_header.csv,nodes_part_*.csv \
  --relationships=rels_header.csv,rels_part_*.csv \
  neo4j

Speed vs Cost: Start with the Free Levers

My instinct when imports were running slow was to throw better hardware at the problem – faster SSDs, locally-attached NVMe. And to be fair, storage does matter. I’ve seen significant speed improvements moving from network-attached storage to high-performance block storage.

What I wish I’d done first, though, is tune the options that cost nothing.

Optimisation Priority – Free First, Then Paid 1 --threads Match to CPU cores Reduce contention FREE 2 --max-off-heap-memory Fewer disk merge passes Biggest single win FREE 3 --high-parallel-io For fast storage only More concurrent ops FREE ▼ Only if the above isn't enough ▼ 4 Upgrade Storage Hardware NVMe / locally-attached SSD / high-perf block storage $$$
Tune the free levers first – storage upgrades should be your last resort, not your first.

The --threads option was the biggest surprise for me. On a machine with plenty of cores, increasing the thread count made a noticeable difference. Conversely, on a smaller instance where I was competing for resources, dialling threads down actually helped by reducing contention.

The --max-off-heap-memory lever was the single most effective change I made. The importer uses off-heap memory for sorting, and if you’re not allocating enough, it falls back to disk-based merge passes – which is slow. On machines with available RAM, bumping this up is a no-brainer.

neo4j-admin database import full \
  --overwrite-destination \
  --threads=16 \
  --max-off-heap-memory=16G \
  --high-parallel-io=true \
  --nodes=nodes_header.csv,nodes_part_*.csv \
  --relationships=rels_header.csv,rels_part_*.csv \
  neo4j

Quality vs Speed: The —bad-tolerance Balancing Act

Getting --bad-tolerance right took me a few painful iterations. Set it too low and every import run dies after a handful of errors – fix, re-run, repeat. Set it too high and you’re drowning in a log file full of noise where it’s hard to find the critical issues.

The --bad-tolerance Iteration Workflow Start with moderate tolerance --bad-tolerance=1000 Run import → review errors Categorise by root cause Critical errors remaining? YES Fix CSV pipeline → re-run NO Only acceptable errors remain Final import config: --bad-tolerance=999999 --skip-bad-entries-logging=true Suppress known errors → fast import
Iterate: start moderate, fix root causes, then open up tolerance for the final clean run.

What worked for me was starting with a moderate tolerance, reviewing the error report to understand the categories of issues, fixing root causes in my CSV pipeline, and then gradually tightening tolerance. The optimal number depends entirely on your data.

Eventually you reach a point where the remaining errors are acceptable – for example, relationships pointing to nodes that belong to a future import iteration. At this point, crank tolerance up and suppress logging:

neo4j-admin database import full \
  --overwrite-destination \
  --threads=16 \
  --max-off-heap-memory=16G \
  --bad-tolerance=999999 \
  --skip-bad-entries-logging=true \
  --auto-skip-subsequent-headers=true \
  --nodes=nodes_header.csv,nodes_part_*.csv \
  --relationships=rels_header.csv,rels_part_*.csv \
  neo4j

Tip: If you’re loading from multiple CSV files that share a schema, enable --auto-skip-subsequent-headers. Without it, redundant header rows in subsequent files get treated as bad data, inflating your error count and sending you on a wild goose chase.

Keep Data Cleansing Out of the Import

This is an opinion I feel strongly about. The import tool offers options like --ignore-empty-strings, --trim-strings, and --normalize-types that handle basic data cleansing during load. They’re tempting – but in my experience, not worth it.

Separation of Concerns: Cleansing vs Import ✗ Avoid: Mixed Responsibilities Raw Data → ETL Pipeline neo4j-admin import --trim-strings=true --ignore-empty-strings=true --normalize-types=true Slower • Harder to debug ✓ Better: Clean Separation Raw Data → ETL Pipeline Trim, validate, normalise here Output: clean CSVs neo4j-admin import Just loads – fast & predictable
Cleanse upstream in your ETL. Let the importer focus on one job: loading data fast.

I once had --normalize-types silently coerce values in ways I didn’t expect, and debugging that was painful because I wasn’t looking at the import step as the source of the problem. When cleansing and import are mixed, you end up unsure whether a data issue came from your transformation code or from an import option you forgot you’d enabled.

Clean your data upstream. Let the importer just import.

Watching the Clock on Large Imports

When you’re importing billions of records, the process can run for hours. I’ve found it helpful to redirect console output to a log file so I can review progress stages and throughput rates after the fact.

# Redirect output while still watching in real-time
neo4j-admin database import full \
  --overwrite-destination \
  --threads=16 \
  --max-off-heap-memory=16G \
  --nodes=nodes_header.csv,nodes_part_*.csv \
  --relationships=rels_header.csv,rels_part_*.csv \
  neo4j 2>&1 | tee import_$(date +%Y%m%d_%H%M).log

It also helps to partition input files into manageable chunks – not just for performance, but for debuggability. When something goes wrong, it’s much easier to trace the issue to a specific file rather than hunting through a monolithic CSV.

Wrapping Up

The bulk import tool is powerful, but the defaults are tuned for convenience, not for large-scale performance. The biggest wins I’ve found come from tuning threads and memory (free), being strategic about --bad-tolerance (saves time), and keeping data cleansing separate from import (saves sanity). Faster storage helps too, but it should be your last resort, not your first.

For the full list of available options and their current defaults, refer to the official Neo4j documentation. Option names and behaviours can change between versions, so always cross-reference with the docs for your specific Neo4j release.

Saurabh Mehrotra

Director at Xpergia

Part of the Xpergia team helping enterprises transform through practical AI implementation.

Explore other Articles

Technical

From RAG to Agents: Building a Grounded Assistant on Amazon Bedrock

How we built the assistant on this site – retrieval that keeps it honest, a relevance floor that makes it refuse, and one real tool call that turns a conversation into a booked meeting.

July 21, 2026 9 min read
Saurabh Mehrotra Director at Xpergia
Read more
Technical

Model Context Protocol in the Enterprise: What It Solves, and What It Doesn't

MCP standardises how agents reach your tools and data, which removes a real integration tax. It does not solve permissions, auditability, or knowing which tools an agent should have.

August 1, 2026 7 min read
Saurabh Mehrotra Director at Xpergia
Read more
Technical

Generative AI Learning Series: Part 1 - Introduction to Artificial Intelligence

Learn what Artificial Intelligence is, why it became necessary, and how it evolved into Generative AI. Welcome to the first installment of our comprehensive series on Generative AI.

August 13, 2026 10 min read
Vikram K Senior Software Engineer
Read more
Technical

Generative AI Learning Series: Part 2 - Evolution of Artificial Intelligence

Trace the 70-year timeline that led to modern Artificial Intelligence and Generative AI. In Part 1, we established what AI is, cleared up common misconceptions, and defined where Generative AI fits into the grand hierarchy.

August 14, 2026 12 min read
Vikram K Senior Software Engineer
Read more
Technical

Generative AI Learning Series: Part 3 - Understanding Machine Learning

Discover how Machine Learning transforms computing by learning patterns from data, exploring its workflow, paradigms, and interactive simulations. In Part 2, we explored how AI evolved from relying on rigid, handwritten rules (Symbolic AI) to systems that can adapt.

August 18, 2026 12 min read
Vikram K Senior Software Engineer
Read more
Technical

Generative AI Learning Series: Part 4 - Neural Networks Explained

Discover how human biology inspired Deep Learning, and explore the mathematical magic behind artificial neurons and deep networks. In Part 3, we saw how Machine Learning shifted the paradigm from explicitly writing rules to teaching computers via examples.

August 19, 2026 14 min read
Vikram K Senior Software Engineer
Read more
Technical

Generative AI Learning Series: Part 5 - Demystifying the Magic: How Neural Networks Actually Learn

Understand the core mechanics of how modern AI systems actually improve themselves. Imagine giving the same math exam to two students. Student A scores 35/100, while Student B scores 95/100. Student B didn't become better overnight.

August 20, 2026 12 min read
Vikram K Senior Software Engineer
Read more
Technical

Generative AI Learning Series: Part 6 - Why Traditional Neural Networks Were Not Enough

Understand the limitations of early neural networks when dealing with memory, context, and sequential data. So far, we’ve learned how a neural network works. It can identify cats in images, predict house prices, classify spam emails, and recognize handwritten digits.

August 21, 2026 11 min read
Vikram K Senior Software Engineer
Read more
Technical

Generative AI Learning Series: Part 7 - Recurrent Neural Networks (RNNs)

Discover how AI learned to remember the past with Recurrent Neural Networks, unlocking the power of sequential data. "Traditional Neural Networks could recognize patterns, but they had no memory.

August 22, 2026 11 min read
Vikram K Senior Software Engineer
Read more
Technical

Generative AI Learning Series: Part 8 - Long Short-Term Memory (LSTM): Teaching AI What to Remember

Learn how to teach AI what to remember and what to forget using Long Short-Term Memory networks. Welcome back to our Generative AI series! In Part 7, we explored how Recurrent Neural Networks (RNNs) gave AI the gift of memory.

August 24, 2026 13 min read
Vikram K Senior Software Engineer
Read more
Technical

Generative AI Learning Series: Part 9 - Transformers: The Breakthrough That Changed AI Forever

Discover the Transformer architecture, the attention mechanism, and how parallel processing laid the foundation for ChatGPT and modern Generative AI. Welcome back! In [Part 8], we saw how LSTMs gave AI a "smart memory," allowing it to remember important details and forget irrelevant ones.

August 25, 2026 15 min read
Vikram K Senior Software Engineer
Read more
Technical

Generative AI Learning Series: Part 10 - The Complete Transformer Architecture Explained Simply

Discover the inner workings of the Transformer architecture, including Positional Encoding, Encoders, Decoders, and Multi-Head Attention. Welcome back to our beginner-to-advanced Generative AI series!

August 26, 2026 15 min read
Vikram K Senior Software Engineer
Read more
Technical

Generative AI Learning Series: Part 11 - Birth of Generative AI: The Moment AI Started Creating

Discover how Artificial Intelligence transitioned from analyzing data to creating completely new content, and where Generative AI fits in the technology landscape.

August 27, 2026 16 min read
Vikram K Senior Software Engineer
Read more
Technical

Generative AI Learning Series: Part 12 - Large Language Models (LLMs): The Technology Behind ChatGPT, Gemini, and Claude

Understand the core technology powering modern AI assistants, how they learn, and how they generate text. If the Transformer architecture we discussed in Part 10 is the "engine," then a Large Language Model (LLM) is the complete vehicle.

August 28, 2026 14 min read
Vikram K Senior Software Engineer
Read more
Technical

Generative AI Learning Series: Part 13 - Demystifying Prompts, Tokens, Context Windows, Temperature, and Hallucinations

Master the essential inner mechanics of Large Language Models, including prompt engineering, tokenization, context windows, temperature scaling, and hallucinations.

August 31, 2026 15 min read
Vikram K Senior Software Engineer
Read more
Technical

Generative AI Learning Series: Part 14 - Popular Generative AI Models: Understanding What Makes Each Unique

Explore the Generative AI landscape and understand the unique strengths of models like ChatGPT, Gemini, Claude, Midjourney, and more. By this point in the blog series, you've learned: Now it's time to meet the actual AI models that are shaping today's world.

September 1, 2026 12 min read
Vikram K Senior Software Engineer
Read more
Technical

Generative AI Learning Series: Part 15 - Practical Real-World Applications (Part 1)

Discover how Generative AI is transforming healthcare, education, software development, marketing, and everyday life. So far in this series, we've learned what AI is, how it evolved, and the mechanics behind Machine Learning, Deep Learning, Neural Networks, Transformers, and Large Language Models.

September 2, 2026 11 min read
Vikram K Senior Software Engineer
Read more
Technical

Generative AI Learning Series: Part 16 - Practical Real-World Applications (Part 2)

Explore how AI is becoming a universal digital assistant across various professional domains, from lawyers to scientists. In the previous part, we explored how Generative AI is transforming Healthcare, Education, Software Development, Marketing, Customer Support, Finance, Agriculture, Manufacturing,…

September 3, 2026 10 min read
Vikram K Senior Software Engineer
Read more
Technical

Generative AI Learning Series: Part 17 - Prompt Engineering: The Art and Science of Communicating Effectively with AI

Master the most critical skill in the AI era by learning how to craft clear, structured, and effective prompts to get the best possible results from Large Language Models.

September 4, 2026 12 min read
Vikram K Senior Software Engineer
Read more
Technical

Generative AI Learning Series: Part 18 - AI Agents: From Answering Questions to Completing Tasks

Discover the evolution from basic chatbots to autonomous AI Agents that can plan, reason, use tools, and execute complex workflows. So far in this series, we've explored Artificial Intelligence, Machine Learning, Deep Learning, Transformers, Large Language Models, and Prompt Engineering.

September 7, 2026 12 min read
Vikram K Senior Software Engineer
Read more
Technical

Generative AI Learning Series: Part 19 - Challenges and Limitations of Generative AI: Risks, Responsibilities, and Ethical Questions

Explore the risks, ethical challenges, and responsibilities associated with Generative AI, from hallucinations and deepfakes to data privacy. So far, this blog series has focused primarily on the extraordinary capabilities of Generative AI.

September 8, 2026 15 min read
Vikram K Senior Software Engineer
Read more
Technical

Generative AI Learning Series: Part 20 - The Future of Generative AI

Explore where AI is heading and what it means for humanity by diving into Multimodal AI, AGI, ASI, and the future workplace. We have now reached the final part of this series. So far, we've explored: Now let's look ahead. What might AI become over the next decade and beyond?

September 9, 2026 16 min read
Vikram K Senior Software Engineer
Read more
Strategy

What Enterprise AI Agents Actually Are (And What They Aren't)

Everyone is selling AI agents. Very little of what's being sold is an agent. Here's the distinction that decides whether your project delivers or quietly stalls.

July 14, 2026 8 min read
Saurabh Mehrotra Director at Xpergia
Read more
Strategy

Agentic Workflow Automation: Where Agents Beat RPA, and Where They Don't

Rule-based automation is cheaper, faster and more reliable than an AI agent – right up to the point where the input varies. A practical framework for deciding which half of your process belongs to which.

July 28, 2026 7 min read
Saurabh Mehrotra Director at Xpergia
Read more