Skip to content

tensor_parser_aten: tag planned tensor DataPtr with its real device#21289

Open
shoumikhin wants to merge 2 commits into
pytorch:mainfrom
shoumikhin:export-D113384858
Open

tensor_parser_aten: tag planned tensor DataPtr with its real device#21289
shoumikhin wants to merge 2 commits into
pytorch:mainfrom
shoumikhin:export-D113384858

Conversation

@shoumikhin

Copy link
Copy Markdown
Contributor

Summary:
In ATen mode, parseTensor built every planned tensor's storage DataPtr with a
hardcoded c10::DeviceType::CPU, and its options with at::CPU(type). A tensor
whose planned buffer lives on an accelerator (e.g. a CUDA-delegate planned
buffer) therefore got a CPU-tagged data pointer, so the runtime treated device
memory as host memory and the delegate rejected it (Method::init fails in
getDeviceFromPtr on a CPU-tagged device pointer).

Fix: read the device from the serialized tensor's extra_tensor_info
(device_type/device_index, defaulting to CPU when absent, matching the portable
parser), then build the tensor with a single
at::from_blob(ptr, sizes, strides, /storage_offset=/0, deleteNothing,
at::TensorOptions().dtype(type).device(device),
/target_device=/device)
so the storage DataPtr, the TensorImpl device (device(), is_cuda()), and the
dispatch key all agree. Passing target_device makes from_blob skip
getDeviceFromPtr, so the same path works for a real device pointer and for a null
runtime-bound pointer without inspecting the pointer. TRT-only / CPU programs are
unchanged (device defaults to CPU).

Also update internal_set_tensor_data (tensor_util_aten.cpp) to preserve the
source tensor's device instead of hardcoding CPU, so sharing a device input's
storage keeps its device tag.

fbcode and xplat copies kept identical.

Differential Revision: D113384858

Copilot AI review requested due to automatic review settings July 23, 2026 17:46
@pytorch-bot

pytorch-bot Bot commented Jul 23, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21289

Note: Links to docs will display an error until the docs builds have been completed.

⏳ No Failures, 27 Pending

As of commit 4ea0725 with merge base 4a26c64 (image):
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jul 23, 2026
@meta-codesync

meta-codesync Bot commented Jul 23, 2026

Copy link
Copy Markdown
Contributor

@shoumikhin has exported this pull request. If you are a Meta employee, you can view the originating Diff in D113384858.

@github-actions

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes incorrect device tagging for planned tensors in ATen-mode deserialization, ensuring that tensors backed by accelerator memory (e.g., CUDA planned buffers) are constructed with consistent device information across the storage DataPtr, TensorImpl device, and dispatch key. It also prevents CUDA backend delegate initialization from clobbering shared-library temp files when identical partitions are loaded multiple times.

Changes:

  • Read device_type/device_index from extra_tensor_info during ATen tensor parsing and construct tensors via at::from_blob(..., TensorOptions().device(device), /*target_device=*/device) to keep device tagging consistent.
  • Preserve the source tensor’s device when sharing tensor storage in internal::share_tensor_data.
  • Ensure unique temp .so paths in the CUDA backend by adding an atomic counter suffix.

Reviewed changes

Copilot reviewed 3 out of 3 changed files in this pull request and generated 3 comments.

File Description
runtime/executor/tensor_parser_aten.cpp Parse device metadata and rebuild tensors with correct device tagging via from_blob(..., target_device=...).
runtime/core/exec_aten/util/tensor_util_aten.cpp Preserve device when sharing tensor storage data pointers.
backends/cuda/runtime/cuda_backend.cpp Avoid temp .so path collisions when loading identical CUDA partitions by adding a unique counter.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +72 to +76
device_type = raw_device_type == executorch_flatbuffer::DeviceType::CUDA
? c10::DeviceType::CUDA
: c10::DeviceType::CPU;
device_index = static_cast<c10::DeviceIndex>(
s_tensor->extra_tensor_info()->device_index());

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FlatBuffers byte is int8_t (signed) and the accessor returns int8_t, so a stored value >127 is already negative on the wire rather than produced by the cast. I've added an explicit check that rejects a negative accelerator device_index (and thus the -1 "current device" placeholder) from the untrusted PTE before it is used.

Comment thread runtime/executor/tensor_parser_aten.cpp Outdated
Comment on lines +133 to +137
if (s_tensor->shape_dynamism() ==
executorch_flatbuffer::TensorShapeDynamism::DYNAMIC_UNBOUND) {
// Provide fully dynamic tensors with an allocator so they can be resized
// within aten kernels.
// Fully dynamic tensors get an allocator so aten kernels can resize them.
// Device-delegate planned buffers are statically bounded, so a device tensor
// never reaches this CPU-tagged path.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The DYNAMIC_UNBOUND path stays CPU-tagged by design: device-delegate planned buffers are statically bounded (not DYNAMIC_UNBOUND), so a device-annotated tensor does not reach here. Hard-rejecting a non-CPU DYNAMIC_UNBOUND tensor is a reasonable additional guard; I'll track it as a follow-up rather than expand this change, since it cannot occur for the delegate path this PR targets.

Comment on lines 134 to +138
ET_CHECK_OR_RETURN_ERROR(
t_src.mutable_data_ptr() != nullptr,
InvalidArgument,
"Source tensor should have data_ptr not being nullptr.");
// Assign the dataptr as the input tensor dataptr
// Preserve the source device; hardcoding CPU would mis-tag a device input's

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed this is a pre-existing asymmetry with the portable implementation for empty tensors (nbytes()==0, valid nullptr data). It's outside the device-tagging scope of this change; tracking as a follow-up to allow nullptr when nbytes()==0.

Copilot AI review requested due to automatic review settings July 23, 2026 18:56

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 3 out of 3 changed files in this pull request and generated 1 comment.

Comment on lines +169 to +173
// Rebuild so storage DataPtr, TensorImpl device, and dispatch key agree.
// target_device makes from_blob skip getDeviceFromPtr, so the same path
// works for a real pointer and for a null runtime-bound one.
tensor = at::from_blob(
data_ptr.get(),

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Correct that the null-placeholder path relies on target_device making from_blob skip getDeviceFromPtr. An ATen-mode device-parser unit test (building a flatbuffer tensor with extra_tensor_info.device_type=CUDA and asserting device/dtype, plus the unindexed-CPU default) is planned as a dedicated follow-up wired into OSS CMake + CI, mirroring the portable tensor_parser_device_test.

shoumikhin added a commit to shoumikhin/executorch that referenced this pull request Jul 23, 2026
…ytorch#21289)

Summary:

In ATen mode, parseTensor built every planned tensor's storage DataPtr with a
hardcoded c10::DeviceType::CPU, and its options with at::CPU(type). A tensor
whose planned buffer lives on an accelerator (e.g. a CUDA-delegate planned
buffer) therefore got a CPU-tagged data pointer, so the runtime treated device
memory as host memory and the delegate rejected it (Method::init fails in
getDeviceFromPtr on a CPU-tagged device pointer).

Fix: read the device from the serialized tensor's extra_tensor_info
(device_type/device_index, defaulting to CPU when absent, matching the portable
parser), then build the tensor with a single
  at::from_blob(ptr, sizes, strides, /*storage_offset=*/0, deleteNothing,
                at::TensorOptions().dtype(type).device(device),
                /*target_device=*/device)
so the storage DataPtr, the TensorImpl device (device(), is_cuda()), and the
dispatch key all agree. Passing target_device makes from_blob skip
getDeviceFromPtr, so the same path works for a real device pointer and for a null
runtime-bound pointer without inspecting the pointer. TRT-only / CPU programs are
unchanged (device defaults to CPU).

Also update internal_set_tensor_data (tensor_util_aten.cpp) to preserve the
source tensor's device instead of hardcoding CPU, so sharing a device input's
storage keeps its device tag.

fbcode and xplat copies kept identical.

Differential Revision: D113384858
shoumikhin added a commit to shoumikhin/executorch that referenced this pull request Jul 23, 2026
…ytorch#21289)

Summary:

In ATen mode, parseTensor built every planned tensor's storage DataPtr with a
hardcoded c10::DeviceType::CPU, and its options with at::CPU(type). A tensor
whose planned buffer lives on an accelerator (e.g. a CUDA-delegate planned
buffer) therefore got a CPU-tagged data pointer, so the runtime treated device
memory as host memory and the delegate rejected it (Method::init fails in
getDeviceFromPtr on a CPU-tagged device pointer).

Fix: read the device from the serialized tensor's extra_tensor_info
(device_type/device_index, defaulting to CPU when absent, matching the portable
parser), then build the tensor with a single
  at::from_blob(ptr, sizes, strides, /*storage_offset=*/0, deleteNothing,
                at::TensorOptions().dtype(type).device(device),
                /*target_device=*/device)
so the storage DataPtr, the TensorImpl device (device(), is_cuda()), and the
dispatch key all agree. Passing target_device makes from_blob skip
getDeviceFromPtr, so the same path works for a real device pointer and for a null
runtime-bound pointer without inspecting the pointer. TRT-only / CPU programs are
unchanged (device defaults to CPU).

Also update internal_set_tensor_data (tensor_util_aten.cpp) to preserve the
source tensor's device instead of hardcoding CPU, so sharing a device input's
storage keeps its device tag.

fbcode and xplat copies kept identical.

Differential Revision: D113384858
Copilot AI review requested due to automatic review settings July 23, 2026 20:12
shoumikhin added a commit to shoumikhin/executorch that referenced this pull request Jul 23, 2026
…ytorch#21289)

Summary:

In ATen mode, parseTensor built every planned tensor's storage DataPtr with a
hardcoded c10::DeviceType::CPU, and its options with at::CPU(type). A tensor
whose planned buffer lives on an accelerator (e.g. a CUDA-delegate planned
buffer) therefore got a CPU-tagged data pointer, so the runtime treated device
memory as host memory and the delegate rejected it (Method::init fails in
getDeviceFromPtr on a CPU-tagged device pointer).

Fix: read the device from the serialized tensor's extra_tensor_info
(device_type/device_index, defaulting to CPU when absent, matching the portable
parser), then build the tensor with a single
  at::from_blob(ptr, sizes, strides, /*storage_offset=*/0, deleteNothing,
                at::TensorOptions().dtype(type).device(device),
                /*target_device=*/device)
so the storage DataPtr, the TensorImpl device (device(), is_cuda()), and the
dispatch key all agree. Passing target_device makes from_blob skip
getDeviceFromPtr, so the same path works for a real device pointer and for a null
runtime-bound pointer without inspecting the pointer. TRT-only / CPU programs are
unchanged (device defaults to CPU).

Also update internal_set_tensor_data (tensor_util_aten.cpp) to preserve the
source tensor's device instead of hardcoding CPU, so sharing a device input's
storage keeps its device tag.

fbcode and xplat copies kept identical.

Differential Revision: D113384858
shoumikhin added a commit to shoumikhin/executorch that referenced this pull request Jul 23, 2026
…ytorch#21289)

Summary:

In ATen mode, parseTensor built every planned tensor's storage DataPtr with a
hardcoded c10::DeviceType::CPU, and its options with at::CPU(type). A tensor
whose planned buffer lives on an accelerator (e.g. a CUDA-delegate planned
buffer) therefore got a CPU-tagged data pointer, so the runtime treated device
memory as host memory and the delegate rejected it (Method::init fails in
getDeviceFromPtr on a CPU-tagged device pointer).

Fix: read the device from the serialized tensor's extra_tensor_info
(device_type/device_index, defaulting to CPU when absent, matching the portable
parser), then build the tensor with a single
  at::from_blob(ptr, sizes, strides, /*storage_offset=*/0, deleteNothing,
                at::TensorOptions().dtype(type).device(device),
                /*target_device=*/device)
so the storage DataPtr, the TensorImpl device (device(), is_cuda()), and the
dispatch key all agree. Passing target_device makes from_blob skip
getDeviceFromPtr, so the same path works for a real device pointer and for a null
runtime-bound pointer without inspecting the pointer. TRT-only / CPU programs are
unchanged (device defaults to CPU).

Also update internal_set_tensor_data (tensor_util_aten.cpp) to preserve the
source tensor's device instead of hardcoding CPU, so sharing a device input's
storage keeps its device tag.

fbcode and xplat copies kept identical.

Differential Revision: D113384858

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 3 out of 3 changed files in this pull request and generated 1 comment.

Comment on lines +339 to +342
filesystem::path so_path = temp_dir /
(so_blob_key + to_string(get_process_id()) + "_" +
to_string(so_file_counter.fetch_add(1, std::memory_order_relaxed)) +
".so");

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed. Since so_blob_key can be an untrusted, variable-length payload key, it's no longer interpolated into the path at all — the filename is now a fixed executorch_cuda_<pid>_<counter>.so prefix under temp_directory_path(), so a key containing / or ../ can no longer influence the write location. The key is still used only to fetch the blob from the named-data map.

shoumikhin added a commit to shoumikhin/executorch that referenced this pull request Jul 23, 2026
…ytorch#21289)

Summary:

In ATen mode, parseTensor built every planned tensor's storage DataPtr with a
hardcoded c10::DeviceType::CPU, and its options with at::CPU(type). A tensor
whose planned buffer lives on an accelerator (e.g. a CUDA-delegate planned
buffer) therefore got a CPU-tagged data pointer, so the runtime treated device
memory as host memory and the delegate rejected it (Method::init fails in
getDeviceFromPtr on a CPU-tagged device pointer).

Fix: read the device from the serialized tensor's extra_tensor_info
(device_type/device_index, defaulting to CPU when absent, matching the portable
parser), then build the tensor with a single
  at::from_blob(ptr, sizes, strides, /*storage_offset=*/0, deleteNothing,
                at::TensorOptions().dtype(type).device(device),
                /*target_device=*/device)
so the storage DataPtr, the TensorImpl device (device(), is_cuda()), and the
dispatch key all agree. Passing target_device makes from_blob skip
getDeviceFromPtr, so the same path works for a real device pointer and for a null
runtime-bound pointer without inspecting the pointer. TRT-only / CPU programs are
unchanged (device defaults to CPU).

Also update internal_set_tensor_data (tensor_util_aten.cpp) to preserve the
source tensor's device instead of hardcoding CPU, so sharing a device input's
storage keeps its device tag.

fbcode and xplat copies kept identical.

Differential Revision: D113384858
shoumikhin added a commit to shoumikhin/executorch that referenced this pull request Jul 24, 2026
…ytorch#21289)

Summary:

In ATen mode, parseTensor built every planned tensor's storage DataPtr with a
hardcoded c10::DeviceType::CPU, and its options with at::CPU(type). A tensor
whose planned buffer lives on an accelerator (e.g. a CUDA-delegate planned
buffer) therefore got a CPU-tagged data pointer, so the runtime treated device
memory as host memory and the delegate rejected it (Method::init fails in
getDeviceFromPtr on a CPU-tagged device pointer).

Fix: read the device from the serialized tensor's extra_tensor_info
(device_type/device_index, defaulting to CPU when absent, matching the portable
parser), then build the tensor with a single
  at::from_blob(ptr, sizes, strides, /*storage_offset=*/0, deleteNothing,
                at::TensorOptions().dtype(type).device(device),
                /*target_device=*/device)
so the storage DataPtr, the TensorImpl device (device(), is_cuda()), and the
dispatch key all agree. Passing target_device makes from_blob skip
getDeviceFromPtr, so the same path works for a real device pointer and for a null
runtime-bound pointer without inspecting the pointer. TRT-only / CPU programs are
unchanged (device defaults to CPU).

Also update internal_set_tensor_data (tensor_util_aten.cpp) to preserve the
source tensor's device instead of hardcoding CPU, so sharing a device input's
storage keeps its device tag.

fbcode and xplat copies kept identical.

Differential Revision: D113384858
Summary:

The CUDA backend writes each compiled AOTInductor library to a temporary file,
named from its content key plus the process id, before loading it. Two identical
CUDA partitions can share the same content key, so both delegates computed the
same path; loading the second library overwrote the first while it was still in
use, and symbol lookup could then crash.

Append a per-process atomic counter to the temporary file name so every delegate
in a process gets a distinct path. The file is still removed when the delegate is
destroyed, as before; only the name changes.

Differential Revision: D113322459
…ytorch#21289)

Summary:

In ATen mode, parseTensor built every planned tensor's storage DataPtr with a
hardcoded c10::DeviceType::CPU, and its options with at::CPU(type). A tensor
whose planned buffer lives on an accelerator (e.g. a CUDA-delegate planned
buffer) therefore got a CPU-tagged data pointer, so the runtime treated device
memory as host memory and the delegate rejected it (Method::init fails in
getDeviceFromPtr on a CPU-tagged device pointer).

Fix: read the device from the serialized tensor's extra_tensor_info
(device_type/device_index, defaulting to CPU when absent, matching the portable
parser), then build the tensor with a single
  at::from_blob(ptr, sizes, strides, /*storage_offset=*/0, deleteNothing,
                at::TensorOptions().dtype(type).device(device),
                /*target_device=*/device)
so the storage DataPtr, the TensorImpl device (device(), is_cuda()), and the
dispatch key all agree. Passing target_device makes from_blob skip
getDeviceFromPtr, so the same path works for a real device pointer and for a null
runtime-bound pointer without inspecting the pointer. TRT-only / CPU programs are
unchanged (device defaults to CPU).

Also update internal_set_tensor_data (tensor_util_aten.cpp) to preserve the
source tensor's device instead of hardcoding CPU, so sharing a device input's
storage keeps its device tag.

fbcode and xplat copies kept identical.

Differential Revision: D113384858
Copilot AI review requested due to automatic review settings July 24, 2026 01:40
shoumikhin added a commit to shoumikhin/executorch that referenced this pull request Jul 24, 2026
…ytorch#21289)

Summary:

In ATen mode, parseTensor built every planned tensor's storage DataPtr with a
hardcoded c10::DeviceType::CPU, and its options with at::CPU(type). A tensor
whose planned buffer lives on an accelerator (e.g. a CUDA-delegate planned
buffer) therefore got a CPU-tagged data pointer, so the runtime treated device
memory as host memory and the delegate rejected it (Method::init fails in
getDeviceFromPtr on a CPU-tagged device pointer).

Fix: read the device from the serialized tensor's extra_tensor_info
(device_type/device_index, defaulting to CPU when absent, matching the portable
parser), then build the tensor with a single
  at::from_blob(ptr, sizes, strides, /*storage_offset=*/0, deleteNothing,
                at::TensorOptions().dtype(type).device(device),
                /*target_device=*/device)
so the storage DataPtr, the TensorImpl device (device(), is_cuda()), and the
dispatch key all agree. Passing target_device makes from_blob skip
getDeviceFromPtr, so the same path works for a real device pointer and for a null
runtime-bound pointer without inspecting the pointer. TRT-only / CPU programs are
unchanged (device defaults to CPU).

Also update internal_set_tensor_data (tensor_util_aten.cpp) to preserve the
source tensor's device instead of hardcoding CPU, so sharing a device input's
storage keeps its device tag.

fbcode and xplat copies kept identical.

Differential Revision: D113384858

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 3 out of 3 changed files in this pull request and generated 1 comment.

Comment on lines +341 to +344
filesystem::path so_path = temp_dir /
("executorch_cuda_" + to_string(get_process_id()) + "_" +
to_string(so_file_counter.fetch_add(1, std::memory_order_relaxed)) +
".so");
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants