Skip to content

OOM crash (heap out of memory) when resuming long session [1.39.2] #783

Description

@Observert

Description

Running cmd --resume <session-id> --yolo on a long-running agent session crashes with a Node.js FATAL ERROR: Ineffective mark-compacts near heap limit Allocation failed - JavaScript heap out of memory. The process aborts (zsh: abort) and the session is lost.

Environment

  • Command Code: 1.39.2
  • Node: v24.19.0 (/Users/shreyashphakadepawar/.nvm/versions/node/v24.19.0/bin/node)
  • OS: macOS 26.6.2 (Build 25G83), arm64
  • Invocation: cmd --resume 06c2423e-70fd-4af7-8416-4d6a2f8d8e87 --yolo
  • Default Node heap limit on this machine: 2240 MB (v8.getHeapStatistics().heap_size_limit)

Crash output

<--- Last few GCs --->
[20041:0x75280c000] 68684349 ms: Mark-Compact 1938.6 (2098.2) -> 1923.8 (2098.0) MB, pooled: 1 MB, 31.92 / 0.04 ms  (average mu = 0.305, current mu = 0.306) task; scavenge might not succeed
[20041:0x75280c000] 68684395 ms: Mark-Compact 1939.8 (2098.0) -> 1922.8 (2097.1) MB, pooled: 1 MB, 32.96 / 0.04 ms  (average mu = 0.300, current mu = 0.296) allocation failure; scavenge might not succeed

FATAL ERROR: Ineffective mark-compacts near heap limit Allocation failed - JavaScript heap out of memory
----- Native stack trace -----
 1: 0x1043219f4 node::OOMErrorHandler(char const*, v8::OOMDetails const&) [...]
 ...
 39: 0x188ef84e4 start [/usr/lib/dyld]
zsh: abort      cmd --resume 06c2423e-70fd-4af7-8416-4d6a2f8d8e87 --yolo

What happened / steps to reproduce

  1. Run a long-lived agent conversation (this one had been active across compaction, with a large session transcript and many tool calls — screenshots, sub-agents, image reads).
  2. Resume that session headlessly: cmd --resume <session-id> --yolo.
  3. The process climbs to ~1.94 GB, GC runs become ineffective (mark-compact can't reclaim), and Node aborts with OOM.

Suspected cause

The resume path likely loads the full session transcript (which after many turns / compaction / large tool outputs — including read_file image data and agent results — grows large) into memory after the shared/global heap is already near the default Node limit. Because the process was launched by node without --max-old-space-size, it dies at the default 2 GB-ish ceiling rather than degrading gracefully. Possible contributors:

  • Session transcript/context payload loaded in full on --resume.
  • Tool outputs (screenshots/images/large read results from sub-agents) retained in memory instead of being streamed or evicted.
  • No explicit heap sizing or soft-cap/memory-pressure handling in the resume path.

Expected behavior

A long session resume should either:

  • not exceed the default heap (stream/evict large payloads), or
  • surface a clear, recoverable error (e.g. "session too large, please compact/start fresh") instead of a hard abort that loses the work, or
  • raise/self-tune the heap limit (e.g. --max-old-space-size) or chunk the resume.

Additional context

This happened on a session that had already been auto-compacted once and contained a large number of tool calls including image reads. It is the first reproducible-ish OOM in ordinary usage — no custom memory settings were in place.

Impact

  • The resumed session crashes and the user loses the working context.
  • Error message only points at the JS engine, not at Command Code — no hint to reduce session size or compact first.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions