AI NewsDev toolsAnnouncement

Cloudflare adds on-demand profiling for Workers and Durable Objects

Cloudflare now lets Workers and Durable Objects produce a CPU or memory flamegraph of code running in production, from the dashboard or the cf CLI, with a chosen duration and version.

AI News

Editorial2 min read

LinkedInX
Cloudflare blog hero for the Workers on-demand profiling announcement

Image: Cloudflare

Why it mattersA Worker team can now point at the exact function burning CPU or holding memory in production, instead of guessing from logs and aggregate metrics about an evicted isolate.

A Worker that is slow or hits its memory limit used to leave its team with logs and aggregate metrics, and no way to see which function was at fault in production. Cloudflare's new on-demand profiler now gives those teams a CPU or memory flamegraph of a running Worker or Durable Object, captured for a chosen duration from the dashboard or the cf CLI.

The feature is available now from the Observability tab on any Worker. The team picks a profile type, a running version with enough traffic, and a duration, then gets back an interactive flamegraph in the dashboard and a downloadable .pprof file for pprof or other tools. The post names three uses Cloudflare has already made of it on its own Workers.

What the flamegraphs found inside Cloudflare

Cloudflare profiled the Worker that implements its R2 binding, which receives heavy traffic, over 50 seconds. In the table view it spotted genericR2JsonReplacer taking over 5% of CPU time, and the function was calling itself as the replacer in JSON.stringify, so a value nested five levels deep was processed five times. The fix was to stop the double walk, which made the function 2.7 times faster. In the same profile it caught a metrics call that was being made twice when storing the first result in a variable would do, and that one duplicate call was 1% of CPU time on its own.

A separate team had a Worker whose P999 memory sat at 133 MB against the 128 MB Worker limit, which caused frequent "Exceeded Memory" evictions. A heap profile opened in pprof showed Prometheus instrumentation code accounting for 66.7% of allocations, which the team believed had been disabled. Only part of it was, and the still-running paths were paying the memory cost without the data ever leaving the Worker. After removing them, P999 memory dropped from 133 MB to 118 MB, giving 10 MB of headroom below the limit, with P50 moving from 70 MB to 54 MB.

How the profiling runs without stopping the Worker

The Workers Runtime holds the isolate lock only around starting and stopping the profiler: it acquires the lock, starts a V8 CPU profiler sampling at one millisecond, then releases the lock so normal requests continue, and reacquires it only to stop and serialise the result. For Durable Objects the request is routed to the specific metal that owns the named actor, so a team can profile one object anywhere in the world. A Worker needs real traffic for a profile to be meaningful, since Cloudflare runs the profiler on the live isolate rather than starting a new one.

Cloudflare writes that the feature has two limits worth stating: profiling has to be started by hand, so a rare misbehaving window can be missed, and the memory profiler only sees allocations that happen inside the capture, so start-up allocations are invisible. The team says continuous profiling is already in development to capture samples automatically. If a Worker is written in TypeScript, source maps need to be enabled on the project or the flamegraph shows obfuscated function names.

Source

Cloudflare, by Dominik Picheta.

This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.

Share
LinkedInX
Start a project