Runtimes Advanced

Node.js PM2 Memory Leak & Infinite Restart Loop Troubleshooting on CentOS Stream / Rocky Linux

Resolve critical Node.js PM2 infinite restart loops caused by memory leaks on CentOS Stream & Rocky Linux. Diagnose, optimize, and stabilize your application.

πŸ‘¨β€πŸ’»
Senior Systems Architect • Verified in Staging Labs

Resolve critical Node.js PM2 infinite restart loops caused by memory leaks on CentOS Stream & Rocky Linux. Diagnose, optimize, and stabilize your application.

Introduction

Encountering an infinite restart loop with your Node.js application managed by PM2 on CentOS Stream or Rocky Linux can be a critical production issue. This typically manifests as your application becoming unresponsive, intermittently unavailable, or serving 50x errors (e.g., 502 Bad Gateway from Nginx). The underlying cause is often a memory leak within the Node.js process, leading to excessive memory consumption. When the process hits its memory limit or a configured threshold, PM2 or the operating system's Out-Of-Memory (OOM) killer intervenes, terminating the process, which PM2 then dutifully restartsβ€”only for the cycle to repeat.

This guide provides a highly technical, step-by-step approach to diagnose and resolve such memory leak issues, restoring stability to your Node.js applications.

Symptom & Error Signature

You will typically observe the following:

  1. Application Unavailability: Your web application is slow, unresponsive, or returning HTTP 50x errors (e.g., 502 Bad Gateway if Nginx is used as a reverse proxy).
  2. PM2 Status Loop: Running pm2 status shows your application repeatedly cycling through online -> stopping -> stopped -> starting -> online states in quick succession.
  3. High Memory Usage: System monitoring tools (e.g., top, htop, free -h) indicate high memory consumption, often by your Node.js process(es).
  4. Application Logs: PM2 logs may show generic restart messages, or more specific FATAL ERROR: Ineffective mark-compacts near heap limit Allocation failed - JavaScript heap out of memory messages.
  5. System Logs (OOM Killer): The operating system's kernel logs (dmesg or journalctl) might reveal the OOM killer terminating your Node.js process due to insufficient memory.

Typical pm2 status Output:

β”Œβ”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ id β”‚ name         β”‚ namespace   β”‚ version β”‚ mode    β”‚ pid      β”‚ uptime β”‚ cpu  β”‚ mem       β”‚ user      β”‚ watching β”‚
β”œβ”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ 0  β”‚ my-node-app  β”‚ default     β”‚ 1.0.0   β”‚ cluster β”‚ 45123    β”‚ 0s     β”‚ 0%   β”‚ 1.2 GB    β”‚ nodeuser  β”‚ disabled β”‚
β””β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
PM2: App 'my-node-app' (id: 0) has been stopped by PM2.
PM2: App 'my-node-app' (id: 0) has been started.
PM2: App 'my-node-app' (id: 0) has been stopped by PM2.
PM2: App 'my-node-app' (id: 0) has been started.
... (This message repeats infinitely)

Typical pm2 logs <app_name> Output (showing memory errors):

0|my-node-app | <--- Last few GCs --->
0|my-node-app |
0|my-node-app | [45123:0x55dc5b8b0000] 1799200 ms: Mark-sweep (reduce) 4089.4 (4138.5) -> 4089.0 (4138.5) MB, 829.7 / 0.0 ms  (average mu = 0.176, current mu = 0.066) allocation failure;
0|my-node-app | [45123:0x55dc5b8b0000] 1799200 ms: Scavenge (reduce) 4089.4 (4138.5) -> 4089.0 (4138.5) MB, 829.7 / 0.0 ms  (average mu = 0.176, current mu = 0.066) allocation failure;
0|my-node-app | [45123:0x55dc5b8b0000] 1799200 ms: Mark-sweep (reduce) 4089.4 (4138.5) -> 4089.0 (4138.5) MB, 829.7 / 0.0 ms  (average mu = 0.176, current mu = 0.066) allocation failure;
0|my-node-app | [45123:0x55dc5b8b0000] 1799200 ms: Scavenge (reduce) 4089.4 (4138.5) -> 4089.0 (4138.5) MB, 829.7 / 0.0 ms  (average mu = 0.176, current mu = 0.066) allocation failure;
0|my-node-app |
0|my-node-app | <--- JS stacktrace --->
0|my-node-app |
0|my-node-app | FATAL ERROR: Ineffective mark-compacts near heap limit Allocation failed - JavaScript heap out of memory
0|my-node-app |
0|my-node-app | # FailureMessage:
0|my-node-app | ----- Error Stack Trace -----
0|my-node-app |
0|my-node-app | <NMT details>
0|my-node-app |
0|my-node-app | Received signal 6 SIGABRT 45123:0x55dc5b8b0000

Typical dmesg | grep -i oom Output:

[123456.789012] Out of memory: Killed process 45123 (node) total-vm:4321000kB, anon-rss:4100000kB, file-rss:0kB, shmem-rss:0kB
[123456.789012] oom_reaper: `node` killed by OOM killer.

Root Cause Analysis

The core of this problem lies in uncontrolled memory growth. Several factors can contribute to a Node.js application memory leak:

  1. Application-Level Memory Leak: This is the most common cause. The Node.js application code itself holds onto references to objects that are no longer needed, preventing the V8 garbage collector from reclaiming their memory. Common patterns include:

    • Unbound Caches or Arrays: Data structures that grow indefinitely without proper eviction policies.
    • Event Emitters: Listeners added but never removed (.off()), leading to objects being held in memory even after they should be garbage collected.
    • Closures: Functions retaining references to variables in their outer scope that would otherwise be eligible for garbage collection.
    • Global Variables: Accidentally assigning large objects to global scope or constantly adding to global arrays/objects.
    • Third-Party Libraries: Bugs or inefficient memory management within external npm packages.
    • Database/Network Connection Leaks: Connections or query results not properly closed or released.
  2. Insufficient System Resources: The server might simply not have enough physical RAM to handle the application's normal operating memory footprint, especially under load. This isn't a "leak" per se, but it leads to the same symptoms.

  3. Incorrect PM2 Configuration:

    • max_memory_restart: PM2 is configured to restart the application when it reaches a certain memory threshold. If a severe memory leak is present, this threshold is repeatedly hit, causing the infinite loop. While useful for preventing catastrophic system crashes, it doesn't solve the underlying leak.
    • Lack of max_restarts or restart_delay: Without these, PM2 will aggressively restart, exacerbating the instability.
  4. Node.js V8 Heap Limits: Node.js applications run within the V8 JavaScript engine, which has a default memory limit for the "old space" (where long-lived objects reside). For 64-bit systems, this is typically around 1.4 GB (Node.js 12+) by default. If your application legitimately requires more memory, hitting this limit can cause an "out of heap memory" error, even if physical RAM is available.

  5. External Factors: High traffic spikes, denial-of-service attacks, or misconfigured reverse proxies can indirectly contribute by stressing the application and exposing latent memory issues.

Step-by-Step Resolution

Follow these steps to diagnose and resolve the memory leak and stabilize your PM2-managed Node.js application.

1. Confirm the PM2 Restart Loop and Memory Usage

First, gather initial diagnostics to confirm the issue.

# Check PM2 application status
pm2 status

# Tail PM2 logs for your application (replace 'my-node-app' with your app name/ID)
pm2 logs my-node-app --lines 200

# Monitor system-wide memory usage in real-time
htop # or 'top' if htop is not installed (sudo dnf install htop)
free -h

# Check kernel logs for Out-Of-Memory (OOM) killer events
sudo dmesg -T | grep -i oom
sudo journalctl -k -p err | grep -i oom

2. Analyze Application Logs and Enable Detailed Logging

Dive deeper into your application's specific output. Look for any errors or warnings before the restarts. If your application's logging is sparse, temporarily increase its verbosity.

Consider writing logs to dedicated files instead of stdout for easier parsing and persistence, especially in production. PM2 can do this via its configuration.

// Example ecosystem.config.js snippet for dedicated logs
module.exports = {
  apps : [{
    name: "my-node-app",
    script: "app.js",
    // ... other configurations
    log_file: "/var/log/my-node-app/combined.log",
    error_file: "/var/log/my-node-app/error.log",
    out_file: "/var/log/my-node-app/out.log"
  }]
};

Make sure the /var/log/my-node-app directory exists and the user running PM2 (nodeuser or root for system-wide PM2) has write permissions.

sudo mkdir -p /var/log/my-node-app
sudo chown nodeuser:nodeuser /var/log/my-node-app # Adjust user/group as needed
pm2 reload ecosystem.config.js # After updating config

3. Adjust PM2 Configuration for Temporary Stability and Diagnosis

Modify your PM2 configuration to prevent continuous, rapid restarts, which can make debugging harder and consume excessive CPU.

These changes are for diagnosis and temporary stabilization. They do not fix the underlying memory leak. A high max_memory_restart can temporarily prevent PM2 from restarting, giving you more time to inspect a leaking process, but also risks complete system memory exhaustion.

Edit your ecosystem.config.js (or similar PM2 configuration file):

module.exports = {
  apps : [{
    name: "my-node-app",
    script: "app.js",
    exec_mode: "cluster", // Or "fork" if not using cluster mode
    instances: "max",     // Or a specific number, e.g., 2
    max_memory_restart: "1.5G", // Temporarily increase, e.g., to 1.5GB or 2GB
                               // This restarts app if resident memory exceeds this.
    max_restarts: 10,           // Allow a maximum of 10 restarts before stopping completely
    restart_delay: 5000,        // Wait 5 seconds between restarts
    kill_timeout: 10000,        // Give the app 10 seconds to gracefully shut down
    env: {
      NODE_ENV: "production",
      LOG_LEVEL: "debug"        // Temporarily set a higher log level for debugging
    },
    // ... other configurations (log_file, error_file etc. from step 2)
  }]
};

After modifying the configuration, reload PM2:

pm2 reload ecosystem.config.js --env production

4. Diagnose Memory Leaks in the Node.js Application

This is the most critical and often complex step. You need to identify where memory is being consumed.

a. Heap Snapshots with Chrome DevTools (Recommended for Production)

This method allows you to take snapshots of the Node.js heap and analyze them offline using Chrome DevTools.

  1. Install heapdump (or use built-in inspector):

    cd /path/to/your/node-app
    npm install heapdump --save
    

    If using Node.js 10+ you can also use the built-in inspector module, but heapdump provides a simpler API for programmatic snapshots.

  2. Integrate heapdump into your application: Add the following lines early in your application's entry point (e.g., app.js):

    // app.js
    if (process.env.NODE_ENV === 'production' || process.env.NODE_ENV === 'development') {
        const heapdump = require('heapdump');
        const fs = require('fs');
    
        // Ensure a directory for snapshots exists
        const snapshotDir = '/tmp/heap-snapshots';
        if (!fs.existsSync(snapshotDir)) {
            fs.mkdirSync(snapshotDir);
        }
    
        // Example: Trigger a snapshot via an API endpoint or a signal
        // For production, prefer a secure, unexposed endpoint or signal handling
        process.on('SIGUSR2', () => { // Send SIGUSR2 to the process to trigger a snapshot
            const filename = `${snapshotDir}/${Date.now()}.heapsnapshot`;
            console.log(`[HEAPDUMP] Writing heap snapshot to ${filename}`);
            heapdump.writeSnapshot(filename, (err) => {
                if (err) console.error(`[HEAPDUMP] Error writing snapshot: ${err}`);
                else console.log(`[HEAPDUMP] Snapshot written to ${filename}`);
            });
        });
        console.log("Heapdump listener activated. Send SIGUSR2 to process to take snapshot.");
    }
    
    // Your main application logic starts here
    // ...
    
  3. Deploy and Restart:

    pm2 reload my-node-app --update-env # Ensure new env vars are loaded
    
  4. Trigger Snapshots: Monitor your application's memory usage with pm2 monit. Once you see memory steadily climbing, trigger a snapshot. Wait a few minutes (or for another significant memory increase) and trigger a second snapshot.

    To send SIGUSR2 to a PM2 process:

    pm2 reload my-node-app # This command restarts the app, which might not be ideal for capturing a leak.
    # A better approach: Find the PID and send signal directly
    APP_PID=$(pm2 prettylist | grep 'my-node-app' -A 7 | grep 'pid' | awk '{print $NF}')
    if [ ! -z "$APP_PID" ]; then
      kill -SIGUSR2 "$APP_PID"
      echo "Sent SIGUSR2 to PID: $APP_PID"
    else
      echo "Application PID not found. Ensure 'my-node-app' is running."
    fi
    

    Repeat this to get 2-3 snapshots over time as memory grows.

  5. Retrieve and Analyze Snapshots: Download the .heapsnapshot files from /tmp/heap-snapshots to your local machine (e.g., using scp):

    scp user@your_server_ip:/tmp/heap-snapshots/*.heapsnapshot .
    

    Open Chrome (or Edge) DevTools, go to the "Memory" tab, select "Load profile" (the upload icon), and load your .heapsnapshot files. Compare snapshots to identify objects that are growing in count or size between snapshots. Focus on objects with a large "Retained Size" or those that significantly increased from one snapshot to the next.

b. Real-time Monitoring with pm2 monit

PM2 includes a built-in monitoring dashboard that can help you observe memory and CPU usage trends.

pm2 monit

Watch the "Mem" column for your application. If it consistently grows before a restart, you have a leak.

c. Profiling with node --inspect (Local/Development Environment)

While less suitable for direct production debugging, this method is powerful for detailed local analysis.

  1. Stop PM2 and Start with Inspector:

    pm2 stop my-node-app
    # Run your app directly with inspector enabled
    node --inspect app.js
    

    This will output a line like: Debugger listening on ws://127.0.0.1:9229/some-uuid.

  2. Connect Chrome DevTools: Open Chrome, navigate to chrome://inspect. You should see your Node.js target under "Remote Target". Click "inspect" to open DevTools.

  3. Perform Actions and Take Snapshots: Interact with your application (e.g., hit endpoints that you suspect cause leaks). In the DevTools "Memory" tab, take heap snapshots before and after these actions, and compare them.

5. Optimize Node.js Application Code

Once you've identified the source of the leak, refactor your code. Common fixes include:

  • Explicitly Nullify References: Set variables to null when no longer needed, especially for large objects.
  • Remove Event Listeners: Always pair emitter.on() with emitter.off() when the listener is no longer required.
  • Bounded Caches: Use libraries like lru-cache for caches with size limits and eviction policies.
  • Stream Processing: For large data, use Node.js streams to process data chunk-by-chunk rather than loading everything into memory.
  • Update Dependencies: Outdated npm packages might have known memory leaks. Use npm outdated to check and update.
  • Handle Errors Gracefully: Unhandled exceptions can leave resources open or cause unexpected object retention.

6. Configure V8 Heap Limits (If Legitimate High Memory Usage)

If your application legitimately needs more memory than the default V8 heap limit (e.g., for in-memory data processing, large machine learning models), you can increase it. This is a workaround, not a fix for a memory leak.

Increasing the V8 heap limit without addressing a true memory leak will only postpone the inevitable OOM crash and can lead to higher system resource consumption. Only do this if you've confirmed your application requires more heap space and doesn't have a leak.

Edit your ecosystem.config.js:

module.exports = {
  apps : [{
    name: "my-node-app",
    script: "app.js",
    node_args: "--max-old-space-size=4096", // Sets V8 old space to 4GB
    // ... other configurations
  }]
};

Adjust 4096 (MB) to a value appropriate for your application and available server RAM. Reload PM2 after making changes:

pm2 reload ecosystem.config.js --update-env

7. Increase System Resources (Short-Term Mitigation / Last Resort)

If the leak is very subtle, or you need immediate stability while debugging, increasing available RAM or adding swap space can provide a temporary buffer.

a. Add Swap Space (CentOS Stream / Rocky Linux)

Adding a swap file can prevent the OOM killer from terminating processes immediately, but it significantly degrades performance if actively used.

# Create a 4GB swap file (adjust count for desired size, e.g., count=8192 for 8GB)
sudo dd if=/dev/zero of=/swapfile bs=1M count=4096
sudo chmod 600 /swapfile
sudo mkswap /swapfile
sudo swapon /swapfile

# Make swap persistent across reboots
echo '/swapfile none swap sw 0 0' | sudo tee -a /etc/fstab

# Verify swap is active
free -h
b. Upgrade RAM

The most straightforward (but costly) solution for insufficient resources is to provision more RAM for your server or virtual machine.

8. Implement Robust Health Checks and External Monitoring

Beyond PM2's internal mechanisms, integrating external monitoring is crucial for long-term stability.

  • Application Health Endpoints: Implement /health or /status endpoints in your Node.js application that return a 200 OK only if critical components (DB connection, external APIs) are working.

  • Reverse Proxy Health Checks: Configure Nginx (or other reverse proxy) to regularly check these health endpoints and automatically remove unhealthy upstream servers from the load balancing pool.

    # Example Nginx upstream configuration for health checks
    upstream node_backend {
        server 127.0.0.1:3000 max_fails=3 fail_timeout=30s;
        # Enable passive health checks
        zone node_backend_zone 64k;
    }
    
    server {
        listen 80;
        server_name yourdomain.com;
    
        location / {
            proxy_pass http://node_backend;
            proxy_http_version 1.1;
            proxy_set_header Upgrade $http_upgrade;
            proxy_set_header Connection 'upgrade';
            proxy_set_header Host $host;
            proxy_cache_bypass $http_upgrade;
            # Aggressive connection pooling for keep-alive connections
            proxy_next_upstream error timeout http_500 http_502 http_503 http_504;
            proxy_next_upstream_tries 3;
            proxy_connect_timeout 5s;
            proxy_send_timeout 5s;
            proxy_read_timeout 15s;
        }
    }
    
  • Monitoring Alerts: Utilize systems like Prometheus/Grafana, Datadog, or Zabbix to monitor your server's memory usage, Node.js process memory, and PM2 application status. Configure alerts for high memory utilization, repeated process restarts, or HTTP 5xx errors.

By systematically diagnosing the memory leak, optimizing your application code, and fine-tuning your PM2 and system configurations, you can resolve the infinite restart loop and ensure the stability and performance of your Node.js services on CentOS Stream or Rocky Linux.

πŸ‘¨β€πŸ’»

Johnathon Wheeler

Senior Systems Architect & DevOps Engineer • Austin, TX

Connect on LinkedIn →

Johnathon has over 16 years of hands-on experience designing, debugging, and scaling Linux web hosting stacks, container clusters, and high-availability database architectures. Every guide on ButItWorkedLocal is independently tested against Debian 12, Ubuntu 24.04/22.04 LTS, Rocky Linux, and Docker environments to guarantee reproducibility in production.

πŸ›‘οΈ

Our Production Verification Guarantee

Encountering a bug not covered here or running a non-standard kernel configuration? Our solutions are continually refined against real production incidents. Submit an environment trace for our editorial team to replicate.