n8n Self-Hosting: RAM, CPU and Execution Requirements

12 min read

343
n8n Self-Hosting: RAM, CPU and Execution Requirements

n8n Resource Basics

n8n is a workflow automation engine that runs workflows as jobs on a server. Self-hosting means your machine handles workflow execution, HTTP webhooks, queueing, and any file or data processing performed by nodes. Resource needs depend less on “n8n itself” and more on what your workflows do: API calls, data transforms, file handling, and how many workflows run at the same time.

For example, a workflow that polls an endpoint every minute and sends a short message to a chat service uses far less CPU and memory than a workflow that downloads large files, converts them, and stores results. If you enable queue mode and run multiple workers, CPU and RAM scale with the number of concurrent executions you allow. When webhooks are enabled, bursts of incoming requests can create short-lived spikes even if average usage stays moderate.

In practice, you size for concurrency and payload size, then validate with measurements. I’ve seen teams choose a small VM, then hit timeouts because a single workflow step blocks the worker while waiting on a slow upstream API. That pattern looks like “CPU is low but executions pile up,” which is a different problem than raw compute shortage.

Common Sizing Mistakes

People often underestimate how n8n behaves under load because they focus on average CPU usage. n8n executions can wait on I/O, but the worker still holds memory for the execution context, node data, and any buffered payloads. When many executions are active, memory pressure shows up as slower responses, higher latency, and eventually out-of-memory kills.

A second mistake is ignoring queueing and concurrency settings. If you run without a queue and allow many simultaneous executions, the single process can become a bottleneck. If you run with a queue but set worker counts too high for the VM, you trade “slow” for “unstable,” because each worker consumes its own memory footprint.

A third mistake is assuming that RAM requirements are constant. Workflows that pass large JSON objects, base64-encoded files, or multi-megabyte payloads can increase memory usage dramatically. Even if the workflow eventually finishes, the peak memory during parsing and transformation can be the limiting factor, not the steady-state average.

Supporting technologies also matter. If you store n8n data in PostgreSQL, the database server becomes part of the performance picture for workflow history, credentials, and queue state. If you use Redis for queueing, Redis memory and network latency affect how quickly jobs move to workers. Docker adds another layer: container memory limits can cause abrupt failures if you set them too low.

RAM And CPU Planning

Start with a workload inventory: number of workflows, expected execution frequency, average payload size, and the slowest external calls. Then decide your concurrency target: how many executions you want to run at once without backlog. That target maps to worker count and the amount of memory each worker needs.

As a practical baseline, many small self-hosted deployments run comfortably on a few gigabytes of RAM when workflows are light and concurrency is low. The moment you add file processing, higher concurrency, or large payloads, you should expect higher memory needs. CPU needs track the amount of transformation work done inside n8n nodes or custom code nodes, not the number of workflow steps.

One useful approach is to treat n8n like a web service plus a job runner. Webhook handling uses request threads and buffers; job execution uses worker processes and keeps execution state. If you run n8n version 1.x with queue mode enabled, the exact behavior depends on your configuration and the node types you use, so measure after you set your concurrency and worker counts. I once saw a deployment that looked fine at 2 workers, then failed after raising to 6 workers because each worker held large execution data during a file download step.

For CPU, watch for node steps that do CPU-heavy work: image/PDF conversions, compression, encryption, or heavy JavaScript transformations. For RAM, watch for steps that hold large objects in memory: reading files into memory, building large arrays, or storing base64 strings in variables. If your workflows mostly call external APIs and pass small payloads, CPU stays low while latency and queue depth become the limiting factors.

Solutions And Advice

Measure With Real Workloads

Before sizing, run a controlled test that matches your real workflow mix. Use n8n execution logs and your server metrics to capture peak memory and execution duration. In a typical test, you trigger the busiest workflows for 15–30 minutes, then compare peak RAM to the steady-state baseline. If you use Docker, check container memory usage rather than host memory, because the container can hit its limit first.

For metrics, Linux tools like top, htop, and free -m help for quick checks, while Prometheus/Grafana or cloud monitoring gives better time series. If you see memory climbing with execution count, reduce concurrency or change workflow steps that buffer large payloads. If you see CPU spikes during specific nodes, isolate those nodes and move heavy work outside n8n.

As an incidental detail: n8n 1.20.x introduced changes in some runtime behaviors, so always measure on the exact version you plan to run. I’ve also seen teams upgrade and assume performance stays identical, then discover a workflow step now returns larger payloads due to a changed API response.

Set Concurrency And Workers

Configure queue mode and worker count to match your RAM budget. Each worker consumes memory for execution context, node data, and any in-memory payloads. A safe starting point is to set workers low, then increase until you see stable latency and no memory pressure. If you run multiple workers, also confirm that your database and Redis can handle the additional job churn.

When you tune concurrency, focus on the bottleneck. If external APIs are slow, adding workers may increase backlog churn without improving throughput. If your workflows are CPU-heavy, adding workers can improve throughput until CPU saturates. If your workflows are I/O-heavy with small payloads, you may hit database or Redis latency before CPU becomes the limiter.

In practice, you’ll often get better results by limiting concurrency for heavy workflows and allowing higher concurrency for lightweight ones. That requires separating workflows or using queue priorities, depending on your setup.

Control Payload Size

Reduce the amount of data you move through n8n. Avoid passing entire files as base64 strings inside variables when a streaming approach is possible. If a node downloads a large file, store it in object storage and pass a reference (URL or key) to downstream steps. When you must transform data, keep transformations incremental so peak memory stays lower.

Also watch for nodes that expand data structures. A workflow that enriches records by adding many fields can grow JSON objects quickly, which increases memory usage per execution. If you only need a few fields, map them early and drop the rest. That reduces both memory and the time spent serializing data between nodes.

A mild frustration many teams hit: the docs show a simple example with tiny payloads, but their real API returns much larger responses, and memory usage jumps. The fix is to test with representative payload sizes, not sample payloads.

Plan For Database And Queue

n8n stores execution history and configuration in a database. If you use PostgreSQL, ensure it has enough CPU and RAM to handle writes from frequent executions. If you enable queueing with Redis, size Redis for the job volume and monitor memory usage. Queue performance depends on network latency too, so keep Redis and PostgreSQL close to the n8n workers in the same region or network.

For outcomes, aim for stable execution latency under peak load and a queue depth that drains after bursts. If queue depth grows without returning to baseline, you need more workers, faster upstream services, or fewer concurrent executions. If the database shows high write latency, you may need to tune indexes, connection pooling, or retention settings for execution history.

As a practical aside, check retention settings for execution logs and history. Keeping months of history can increase database size and slow queries that n8n runs for UI and reporting.

Case Examples

Webhook Burst With Small Payloads

A small team self-hosts n8n to receive webhook events from an internal system. Each execution calls two external APIs and writes a short record to PostgreSQL. They start with 2 workers on a 4 vCPU VM with 8 GB RAM and observe that CPU stays under 20% while memory peaks around 2.5–3.5 GB during bursts. When they raise workers to 6, memory peaks near the limit and latency increases, even though CPU remains low. The team reduces workers back to 3 and adds rate limiting at the webhook source, which keeps queue depth stable.

File Processing With Large Inputs

A media workflow downloads PDFs, extracts text, and sends summaries to a downstream service. The PDFs average 15–25 MB, and the extraction step returns large text blobs. On a 2 vCPU VM with 4 GB RAM, executions start but the container restarts during peak hours due to memory pressure. After moving the extraction step to an external service and storing extracted text in object storage, n8n executions pass only references and small metadata. With the same VM, peak memory drops and executions complete without restarts, while throughput becomes limited by the external extraction service rather than n8n memory.

Resource Checklist And Table

Parameter What It Affects How To Measure Tuning Direction
Worker Count Parallel executions and memory per execution Peak RAM vs queue drain time Lower if memory spikes; raise if CPU is low and queue grows
Payload Size Execution context memory and serialization time Track node outputs and response sizes Reduce fields, store files externally, pass references
External API Latency Execution duration and queue backlog Execution time breakdown per node Add retries/timeouts, reduce concurrency for slow workflows
Database/Redis Job state, history writes, queue throughput DB write latency and Redis memory Tune retention, indexes, pooling; scale DB/Redis if needed

Step-by-step checklist for sizing a new deployment:

  1. List workflows by type: webhook-driven, scheduled, and manual; record expected frequency.
  2. For each workflow, estimate peak payload size and identify nodes that transform or buffer data.
  3. Choose a target concurrency for heavy workflows and a separate target for lightweight workflows.
  4. Start with low worker count, then run a 15–30 minute load test using representative inputs.
  5. Increase workers until queue depth stabilizes and peak memory stays below your container/VM limit.
  6. Verify database and Redis latency during the test; scale them if they become the bottleneck.
  7. Set alerts for memory pressure, queue depth, and execution failure rates.

Common Mistakes

One mistake is sizing from a “hello world” workflow. A single HTTP request node with tiny payloads hides the memory impact of large outputs and file handling. Another mistake is ignoring container memory limits; Docker can kill the process when it hits the limit even if the host has free RAM.

Teams also misread metrics. Low CPU does not mean the system has spare capacity; it can mean workers are waiting on network calls while executions accumulate. In that situation, raising workers can increase memory usage without improving throughput, which looks like “it got worse after scaling.”

Another practical error is leaving execution history retention too long. Large history tables can slow down UI queries and background tasks, which increases latency and can indirectly affect execution scheduling. If you need history for compliance, store it, but tune retention and indexing so the operational workload stays manageable.

Finally, people sometimes treat timeouts as a performance problem rather than a reliability design choice. If upstream APIs respond slowly, set timeouts and retries so executions fail fast when needed. That keeps workers available for other jobs instead of tying up memory for long waits.

FAQ

How Much RAM Does n8n Need?

RAM depends on workflow concurrency and payload size. Light workflows with low concurrency often fit in a few gigabytes, while file processing and large JSON outputs can raise peak memory per execution. Measure peak container/VM memory during a load test that matches real inputs.

How Many CPU Cores Should I Use?

CPU needs track CPU-heavy nodes and the number of concurrent workers. If workflows mostly wait on external APIs, CPU stays low and queue depth becomes the limiter. If workflows transform large data in-process, CPU saturation appears during those steps.

Do I Need Redis For Self-Hosting?

Redis is commonly used for queueing in self-hosted setups, but the exact requirement depends on your configuration. If you use queue mode, Redis helps manage job flow; if you run without queueing, you still need to manage concurrency to avoid overload.

What Causes Memory Spikes During Executions?

Memory spikes usually come from large payloads held in execution context, base64-encoded file data, or nodes that build large in-memory objects. Slow upstream calls can also increase the number of active executions, which multiplies memory usage across workers.

How Do I Know My Server Is Under-Provisioned?

Under-provisioning shows up as growing queue depth, rising execution latency, and increased failure or timeout rates. Low CPU with increasing backlog often indicates I/O waits or database/Redis latency rather than raw compute shortage.

Author's Insight

n8n resource planning works best when you treat workflows as a mix of web requests and background jobs with measurable concurrency. RAM sizing depends on how much data each execution holds at peak, not on the number of nodes in the workflow. CPU sizing depends on whether nodes perform CPU-heavy transformations inside the n8n process. Queueing and external services shift the bottleneck from compute to latency and storage writes, so monitoring queue depth, database latency, and peak memory together gives a clearer picture.

If you share your workflow types, expected execution frequency, and approximate payload sizes, you can turn this into a concrete sizing plan with a load test and a safe worker count. Without those inputs, any single “RAM/CPU number” stays guesswork.

Key Takeaways

  • Size n8n by concurrency and payload size, then validate with peak memory measurements during a realistic load test.
  • Low CPU does not guarantee capacity; queue depth and database/Redis latency often reveal the real bottleneck.
  • Limit memory growth by reducing payloads and storing large files outside n8n when possible.
  • Tune worker count and queue settings together, and keep database and Redis performance in the same monitoring view.

Was this article helpful?

Your feedback helps us improve our editorial quality

Latest Articles

Automation 23.09.2026

Zapier vs Make: Task Limits and Execution Costs

Zapier and Make both automate work between apps, but their pricing and limits hinge on how many tasks run and how executions are counted. This article helps readers compare task limits, execution costs, and common billing surprises using concrete examples like email-to-CRM sync and form-to-sheet logging. You’ll learn how to estimate monthly runs, spot hidden consumption drivers, and choose a plan that matches real workflows without guessing.

Read » 401
Automation 17.09.2026

Automation Error Handling: Fail Fast vs Retry

Automation error handling decides what a system does after a failure: stop immediately (fail fast) or try again (retry). This article explains how those choices affect reliability, safety, and user trust in automated workflows. It is for engineers, operations teams, and informed readers who want to evaluate automation behavior in real systems. You will learn failure modes, retry design limits, backoff and idempotency, and practical checklists with examples.

Read » 284
Automation 30.08.2026

OAuth vs API Keys: Which Is Safer for Automations?

Learn how OAuth and API keys work in real automation workflows, with a focus on safety: token theft, scope control, rotation, and audit trails. It’s for people building or maintaining integrations for health-related services and other regulated systems. You’ll learn how each method behaves in practice, what to check in provider docs, how to reduce blast radius, and which failure modes to plan for before you ship.

Read » 415
Automation 18.08.2026

Webhooks vs Polling: Which Automation Method Wins?

Webhooks and polling are two ways to automate updates between systems, such as patient portals, lab feeds, and appointment tools. This article explains how each method works, where delays and failures come from, and how to choose based on timing needs, reliability, and cost. Readers will learn practical design checks, common mistakes, and decision criteria using real-world examples and a comparison checklist.

Read » 204
Automation 24.08.2026

API Rate Limits: Why Your Workflow Suddenly Stops

API rate limits can halt a health-related workflow without warning: a script stops syncing data, a dashboard shows stale results, or a form submission fails. This article explains how rate limits work, why they trigger suddenly, and how to diagnose the cause using headers, logs, and retry behavior. It also covers practical fixes like backoff, batching, and quota planning, plus common mistakes that lead to repeated outages.

Read » 330
Automation 11.09.2026

Retry Logic: How Many Times Should a Workflow Retry?

Retry logic controls how a workflow reacts to failures by trying again after a delay. This article explains how many retries to use, how to choose retry delays, and when retries create risk instead of resilience. It is for engineers and health-adjacent teams building or auditing automated workflows that touch patient data, appointments, claims, or lab results. You’ll learn practical retry limits, failure classification, and how to test behavior so systems recover without amplifying outages.

Read » 133