X-Ops

Ray silent door: how CVE-2025-62593 lets attackers run shell commands through your browser

When Anthropic's research team released Ray as an open-source framework for scaling Python and AI workloads, they made an architectural decision that has resurfaced as a critical vulnerability three times in two years: the most powerful endpoints on a Ray cluster would never require authentication. The /api/jobs endpoint, the one used to submit and manage distributed jobs, has remained open since the project's earliest releases. In August 2026, that decision became a federal incident.

CVE-2025-62593 was added to CISA's Known Exploited Vulnerabilities catalog on August 17, 2026, with a remediation due date of August 20, 2026 for federal agencies. The CVSS 3.1 base score is 9.4. The CISA SSVC classification, updated the same day the entry was published, marks exploitation as "active," automatable as "no," and technical impact as "total." It is the third time in twenty-four months that the absence of authentication on Ray's job submission endpoints has produced a critical CVE, and the first time the exploitation vector has shifted from the network to the browser.

The chain, end to end

CVE-2025-62593 is not a memory corruption vulnerability. It is the convergence of three weaknesses that are individually well-understood but have never been combined in quite this way against a developer framework.

Weakness one is the unauthenticated endpoint. Ray's /api/jobs and /api/job_agent/jobs/ endpoints accept job submissions from any client that can reach the dashboard HTTP port. By default this is 8265, but in custom deployments the service ports (6379, 10001) are frequently exposed. The job submission API takes a runtime environment specification; the Raylet on the target node processes it; in versions prior to 2.52.0, that processing path included insufficient validation of which fields were honored when the request came from a browser-like User-Agent.

Weakness two is User-Agent handling. According to the Ray maintainers' own advisory, the issue stems from "insufficient controls against browser-based attacks, specifically scenarios where the User-Agent header can be modified." Ray's code path for job submission did not properly distinguish between a real Ray client, which uses a custom runtime and emits a distinctive fingerprint, and a request originating from a malicious web page running in Firefox or Safari.

Weakness three is DNS rebinding. Modern browsers implement varying degrees of protection against DNS rebinding: Chrome pins and rejects rebinding attempts against localhost, Firefox has weaker controls, Safari historically the weakest. The attack works by serving a hostname that initially resolves to an attacker-controlled server, then changing the DNS response after the browser has cached the original resolution. Once the rebinding succeeds, the malicious page can make requests to 127.0.0.1 on any port, bypassing the same-origin policy that normally prevents a web page from talking to local services.

The chain unfolds like this: a developer running Ray on their workstation visits a malicious page or is served a malicious advertisement. The page initiates a DNS rebinding attack against localhost:8265. Once the rebinding succeeds, the page submits a job to Ray's API with a payload that embeds shell commands in the runtime environment specification. The Ray cluster executes those commands with the privileges of the user running the Ray process.

Who is exposed

The primary victims are developers running Ray in development and testing environments: machine learning engineers iterating on training jobs, data scientists exploring datasets, platform teams running Ray on shared infrastructure for experimentation. Ray's design as a developer-facing framework means production deployments are less common, but the exposure surface in development environments is enormous.

The risk is amplified by the deployment patterns typical for ML workloads. Ray clusters are frequently spun up on cloud instances with the dashboard port exposed for remote access from a laptop. Developers running Ray on their workstation often have source code, secrets in environment variables, SSH keys for production systems, and credentials for cloud APIs in scope. A successful RCE gives an attacker not just the local machine but a pivot into the broader infrastructure the developer can reach.

CI and CD pipelines deserve particular attention. Many teams integrate Ray into training pipelines, often running Ray on long-lived instances that hold credentials for object storage, container registries, and model artifact repositories. The vulnerability does not require the attacker to know the cluster's network location; DNS rebinding works against localhost, and any browser running on the same machine as a Ray process is a viable target.

For cloud-hosted Ray clusters, the situation is more nuanced. If the cluster's dashboard port is exposed via a public IP, a common pattern for letting remote developers connect from laptops, then CVE-2025-62593 is remotely exploitable without the DNS rebinding component. The browser-based attack is the novel twist, but the underlying unauthenticated endpoint has been remotely exploitable on misconfigured clusters for years.

Patching and version specifics

CVE-2025-62593 affects Ray versions prior to 2.52.0. The fix was committed upstream and released as version 2.52.0. The patch addresses User-Agent handling on the affected endpoints and adds additional validation to the runtime environment processing path.

For teams running Ray on Kubernetes via KubeRay, the upgrade depends on the operator version. KubeRay 1.14.0 and later support Ray 2.52.0. Verify the Helm chart or operator deployment allows the image override, then update the Ray image tag. Rolling restarts are supported.

For teams running Ray directly on virtual machines or bare metal, the upgrade is a process restart. The Ray head node and worker processes must be restarted to pick up the new binary. There is no in-place patch.

For environments where an immediate upgrade is not feasible, four mitigations reduce the attack surface without eliminating the vulnerability.

Network-level controls come first. Restrict the Ray dashboard port to known IP ranges. On cloud infrastructure, ensure security groups do not expose 8265, 6379, 10001, or other Ray service ports to the public internet. This is the single highest-impact mitigation and should be standard practice regardless of CVE status.

A reverse proxy with authentication is the second lever. Place a reverse proxy that requires authentication in front of the dashboard. nginx, Envoy, or Caddy can terminate TLS and require a valid bearer token before forwarding to the Ray cluster. This addresses both CVE-2025-62593 and the broader class of unauthenticated endpoint issues.

Disabling browser-based access is the third option. Teams that do not require browser access to the Ray dashboard can mitigate CVE-2025-62593 specifically by ensuring that the dashboard is only accessed through programmatic clients with non-browser User-Agents. This is partial; the underlying unauthenticated endpoint issue persists, but it removes the specific attack vector.

Browser-level DNS rebinding protections are the fourth band-aid. Firefox users can enable network.dns.disableIPv6 and network.dns.disablePrefetch. Safari users should ensure they are running the latest version. These reduce the rebinding window but do not eliminate it.

What this tells us about Ray's security model

The Ray maintainers' advisory is unusually candid about the root cause: "Due to the longstanding decision by the Ray Development team to not implement any sort of authentication on critical endpoints, like the /api/jobs & /api/job_agent/jobs/, has once again led to a severe vulnerability that allows attackers to execute arbitrary code against Ray." The phrase "once again" is significant. CVE-2022-1525, CVE-2023-0706, and CVE-2023-48022 followed the same pattern. Each time the root cause was the same. Each time the fix was endpoint-specific hardening or a stopgap that left the architectural issue intact.

Ray is not unique in making this trade-off. Many developer-facing tools ship with permissive defaults on the assumption that the user is operating in a trusted environment. Kubernetes requires authentication on its API server but defaults service account permissions to broad and has historically had weaker authentication on the kubelet. Docker daemon sockets have been an attack surface since the project's inception. The issue with Ray is not that it made a single bad decision; it is that the decision was institutionalized and defended as a feature rather than recognized as a liability.

For teams evaluating Ray for production use, CVE-2025-62593 is a marker. It is not a reason to avoid Ray, because the framework's capabilities for distributed ML workloads remain unmatched in the open-source ecosystem, but it is a reason to treat any Ray deployment with the same operational discipline applied to any production-critical service. That means authentication at the network boundary, monitoring of the dashboard port for anomalous access, and a documented upgrade path that can be executed within hours when a critical CVE is published.

Detection and hunting

Blue teams looking for signs of exploitation should focus on three telemetry sources.

Ray cluster logs are the first. Job submissions from unexpected sources, particularly those with browser-like User-Agent strings or originating from IP addresses not associated with known clients, are an immediate red flag. Ray's logging captures job submission metadata including the User-Agent, and the legitimate Ray client uses a distinct fingerprint that does not match any mainstream browser.

Process execution telemetry on hosts running Ray is the second source. A successful RCE chain via CVE-2025-62593 results in the Ray process spawning a child process to execute the attacker's payload. On Linux this manifests as a child of the Ray Python process invoking a shell (/bin/sh, /bin/bash) or a network utility (curl, wget, nc). EDR products should flag child processes of the Ray runtime that do not match the expected training-job profile, especially those that appear shortly after an outbound DNS resolution to a recently registered domain.

DNS resolution patterns are the third. DNS rebinding attacks rely on the attacker's DNS server returning a second response with a different IP address than the first. Monitor DNS query patterns for short TTL responses to hostnames that subsequently resolve to RFC1918 addresses or to localhost. Most corporate resolvers log these transitions; the difficulty is correlating them with browser-side activity, which typically requires endpoint telemetry.

The bigger picture

CVE-2025-62593 is a useful case study because it combines multiple classes of vulnerability, including unauthenticated endpoints, browser-based attack surfaces, and DNS rebinding primitives, into a single high-impact RCE. Each component has been studied extensively in isolation. Together, they produce a chain that bypasses the typical defenses developers rely on for local development tools.

The lesson is not new but bears repeating: developer-facing tools that are easy to use locally are also easy to attack locally. The same characteristics that make Ray pleasant for iterating on training jobs (automatic port exposure, permissive defaults, no authentication requirements) make it a target. As more AI and ML infrastructure moves into production environments, the security boundary between local developer tools and production infrastructure will continue to blur. CVE-2025-62593 is a preview of the kind of incident that boundary blurring produces.

For platform teams, the practical takeaway is straightforward: inventory every Ray deployment in your organization, verify the version, apply the upgrade to 2.52.0 or later, and ensure that the dashboard is not exposed beyond the network boundary where the developer is operating. For developers, the lesson is to assume that any tool running on your workstation with network exposure is part of your attack surface, because it is. Treat local services with the same skepticism you would apply to a public-facing production endpoint, and the next CVE in this category becomes a configuration review instead of an incident response.

Compliance and reporting obligations

For organizations subject to regulatory reporting, CVE-2025-62593 carries specific obligations worth noting. Under CISA's Binding Operational Directive 22-01, federal agencies must remediate vulnerabilities listed in the KEV catalog within the prescribed timeframe; the August 20, 2026 due date for this CVE was three days from publication. Organizations operating under PCI-DSS, HIPAA, or SOC 2 frameworks should treat KEV-listed CVEs as priority findings for their own internal patching SLAs, even though those frameworks do not directly reference the KEV catalog. Insurance carriers increasingly require evidence of KEV remediation within the CISA-mandated window as a condition of cyber liability coverage, and failure to patch within that window has been cited in post-incident disputes as evidence of negligence.

If your organization runs Ray and you discover indicators of exploitation, the incident response playbook should treat it as a workstation or CI runner compromise, not a single-host event. The credentials accessible to the user running Ray at the time of exploitation should be considered exposed; secrets in environment variables, SSH keys, cloud credentials, and tokens for source repositories all need rotation. Forensic analysis should focus on the Ray process tree, the user account that owned the Ray process, and any outbound network connections initiated by child processes of the Ray runtime in the days surrounding the suspected compromise window.

The bigger lesson for security teams is to stop treating developer workstations as trusted environments. The Ray incident is one of a growing category in which a vulnerability in a developer tool produces a path into production infrastructure. The same pattern appears in IDE vulnerabilities, in CI runner compromises, and in container escape chains from build hosts. Each of these categories shares a common theme: the developer environment is on the critical path to production, but its security posture is rarely held to the same standard as the production environment itself. CVE-2025-62593 is a reminder that this gap has costs, and those costs are now landing on the same incident response teams that were told the local tools were not their problem.