Skip to content

Fix: kernel-mode capacity refusals no longer free what they protect (⑤K3) - #2193

Draft
sunkaixuan2018 wants to merge 2 commits into
hw-native-sys:mainfrom
sunkaixuan2018:skx/kernel-capacity-freeze-k3
Draft

Fix: kernel-mode capacity refusals no longer free what they protect (⑤K3)#2193
sunkaixuan2018 wants to merge 2 commits into
hw-native-sys:mainfrom
sunkaixuan2018:skx/kernel-capacity-freeze-k3

Conversation

@sunkaixuan2018

@sunkaixuan2018 sunkaixuan2018 commented Sep 11, 2026

Copy link
Copy Markdown
Contributor

Depends on #2064 (①K1) @ 86dd62b4912866a6e96d63891b4c2838a2635b5a

This branch is built directly on that commit, which is not yet merged. The
main base above is a GitHub constraint — a branch on a fork cannot be a PR
base — so the diff shown against main includes all of K1. Exactly one
commit is mine:

  • b8e739d9 — kernel-mode capacity refusals no longer free what they protect

Incremental diff (this PR only):
sunkaixuan2018/simpler-PTO@86dd62b...skx/kernel-capacity-freeze-k3

Draft until #2064 merges, then rebased and marked ready.

What this is

A kernel-mode context's device buffers must keep their addresses for the life of
the context, because an ACLGraph replays the addresses that were captured. Every
runtime-time growth is release + reserve or free + malloc, which re-bases; the
failure mode is silently wrong data on replay, not a crash. That makes this a
correctness change, not a sizing one.

K1 landed the first guard on that invariant, in setup_static_arena. Its
predicate is right and is unchanged here.
What this PR fixes is that its
refusal, and the refusal on the second growth point, were both destructive.

Two refusals that destroyed what they protected

1. setup_static_arena — the guard freed the bases it had just declined to
move.
The refusal returned PTO_RUNTIME_ERR_INTERNAL, the caller collapsed that
to ok = false, and the !ok block released all three regions unconditionally.
DeviceArena::release() frees the backing buffer, so a fired guard dropped
exactly the base addresses it names in its own error message. K1's comment above
the guard said so; nothing acted on it.

2. RetainedTempBump::begin — the free came before anything could refuse. The
grow path is device_free(old) then device_malloc(bigger). A refusal placed
anywhere downstream of that — including at the platform's device_malloc
arrives after the address is already gone, and the existing failure path then
clears the slot to {nullptr, 0}. The guard has to sit ahead of the free.

The invariant both fixes share, and the one every later growth-point guard will
need: a refusal must have no side effects.

What changed

host/static_arena_bank.h (new) the bank's commit rule, shared by the onboard and simulation runners instead of hand-duplicated in both
the rule's failure handling an allocation failure still rolls the whole bank back; a capacity refusal leaves every region committed and every cached size intact
both setup_static_arena bodies ~55 lines each become a request literal plus a call; each keeps its own prebuilt-arena cache invalidation, now keyed on bases_changed
RetainedTempBump::begin (a2a3 + a5) refuses a kernel-mode re-base ahead of the device_free, leaving the slot as the previous run left it
HostApiOps::is_kernel_mode (new) how the context's identity reaches runtime code; a table that does not supply it reports program mode
six contract sites the docs that described the behavior these guards make conditional, listed below

is_kernel_mode is a query of the existing ExecutionModeLatch, not a new
switch — no environment variable and no macro, per
env-macro-gating.md §1. Its producer
(c_api_shared.cpp) and its consumer (runtime_maker.cpp) are compiled into the
same libhost_runtime.so, so appending to the ops table creates no ABI skew.

The retained-buffer guard refuses a re-base, not an allocation

This is the part worth a reviewer's attention, because I got it wrong first.

The guard fires on required > size && addr != nullptr. The addr != nullptr
half matters: an empty slot holds no address a captured graph could reference, so
a context's first allocation moves nothing and is taken normally. Without that
clause the guard refuses the first bind too, and since the grow path is the only
writer of the slot, the slot would stay {nullptr, 0} forever — a kernel context
could never reach its frozen size at all.

My first version omitted it, on the premise that a kernel run never stages host
tensors anyway. That premise is not enforced anywhere in the tree, and I could
not support it on a re-read:

  • validate_kernel_launch_args checks null-ness and the callable-id range. It
    never inspects tensors.
  • ChipTensor::address_space defaults to HOST, so the default goes the
    opposite way.
  • SimplerKernelInvocationHeader::host_copy_tensor_count exists precisely to
    carry host-memory args in kernel mode, documented as "zero until the host-only
    copy contract lands". Host tensors there are a planned contract, not an excluded
    one.

So the narrow guard is the correct one: it still refuses every real re-base, and
it refuses nothing else.

Doc updates in the same commit

The behavior these guards make conditional was described in six places, all of
which now read false for a kernel context and move with the code:
setup_static_arena and the retained-temp-buffer entries in host_api.h,
DeviceRunnerBase::setup_static_arena's rollback promise (and its @return,
which named -1 for a function that returns PTO_RUNTIME_ERR_INTERNAL), the
kernel-mode capacity paragraph in runtime_c_api.h, both runners' latch comments,
docs/task-flow.md, and RUNTIME_LOGIC.md in both architecture trees.

The extraction is a flagged deviation from "program path byte-identical"

Behavior in program mode is unchanged — kernel_mode is false there, so the new
branch is unreachable and growth, release and rollback all take the paths they
took before. But the code moved, so this is not literally untouched, and it is
deliberate for two reasons. The rule was two hand-maintained copies that
codestyle.md §10 warns about, and onboard/sim
symmetry is now structural rather than a review obligation. And it is what makes
the fix testable without a device: DeviceRunnerBase needs CANN, the rule needs
nothing.

The allocate_tensor design question, answered

The handover asked whether kernel mode should refuse allocate_tensor
unconditionally, or only outside a prepare window borrowed from ④K2. The
unconditional form holds
: #2176's prepare-once allocator reaches mem_alloc_
directly through its own context ops, never through HostApi::device_malloc, so
refusing the runtime-facing surface would not block a legitimate kernel-time
allocation. This PR does not install that refusal, because the growth point it is
in scope for needs its guard higher up anyway — but the question is settled and the
next guard does not have to re-derive it.

Out of scope, and one question for a reviewer

Growth points #1 acquire_graph_definition_block and #2 HBG host-tensor staging
are not closed here.
Both are HBG-only and both live in files #2173 rewrites;
#2173 already carries a disposition for the definition block. @TaoZQY — do you want
to keep #1 (and #2) on your side, or should a follow-up here close both against the
is_kernel_mode query this PR adds? Splitting one growth point across two PRs is
the outcome worth avoiding. Note that acquire_graph_definition_block is
grow-by-replacement of a device block, so it is the same silent-replay hazard,
not a lesser one.

DFX is outside the frozen capacity, but it is not address-stable. This is the
question the handover asked, so here is the full answer rather than half of it. The
DFX device buffers — the device-wall buffer and the collectors' workspaces — are
separate mem_alloc_ allocations, not regions of an arena bank, so nothing this PR
freezes covers them and it correctly writes no DFX teardown logic. The device-wall
buffer is allocated lazily once and never grown, so it does not re-base. The
collector pools do
: prepare_execution calls finalize_collectors() whenever
collector_shape_is_stale(...), so a run whose core or AICPU-thread counts differ
from the previous one tears their device memory down and rebuilds it at a new
address. A kernel context's shape is context-static, so that should not fire — but
nothing enforces it, and if the DFX buffers are ever reachable from a captured
graph this is a growth point in its own right. Flagging it rather than guarding it
here: it belongs with whoever owns DFX under the execution claim (#2163).

The sim-side guard stays. Simulation never latches KERNEL, so that branch is
never true, but removing it would break the onboard/sim symmetry the stub-parity
argument rests on.

Verification

All on myserver (aarch64, CANN 9.0.0). No device is needed for any of it.

Lane Result
cpput, no hardware 144/144 passed — base 86dd62b4 is 143/143, built and run the same way
new test_static_arena_bank 11/11, no_hardware label
test_trb_runtime_temp_buffer 16/16, was 10 — and identically for test_a5_trb_runtime_temp_buffer
editable install all 8 libhost_runtime.so built (a2a3 + a5 × onboard + sim × hbg + trb); no warning names a changed file
tests/ut/py/test_host_runtime_abi.py 4 passed, 4 skipped — the exported symbol set is unchanged
pre-commit all hooks pass, clang-tidy included, in CI

The five acceptance criteria are covered twice, once per slot kind:

arena bank retained temp buffer
base and capacity constant across N calls RepeatedCommitsMoveNoBaseAndCallNoAllocator KernelModeHoldsOneAddressAcrossRepeatedRuns
zero underlying alloc/free same, via DeviceArena::alloc_count() same, via the fake's malloc/free counters
exact capacity accepted ExactAndSmallerRequestsAreServedInPlace same test, a run that exactly fills the buffer
over capacity refused with INTERNAL OneByteOverCapacityIsRefusedAndChangesNothing KernelModeRefusesToGrowTheRetainedBuffer
refusal does not break the held plan same test's second half KernelModeRefusalLeavesTheRetainedBufferUsable

KernelModeAllocatesOnceThenRefusesToGrow is the production shape: the latch is
write-once, so a real context is in kernel mode from its very first bind. It
asserts one allocation, then a refusal, then the sized run still binding from the
same address.

The other exit is pinned too. AllocationFailureRollsTheWholeBankBack and
AllocationFailureOnTheFirstCommitStillRollsBack drive a failing backing
allocator through DeviceArena's injectable alloc/free, so the rollback path is
reachable with no device: an allocation failure releases every region and zeroes
every remembered size, in kernel mode exactly as in program mode, because a setup
that never completed published no address to protect.

Negative controls

A test that passes against the fix is not evidence it would have caught the
defect, so each defect was re-introduced and the suites re-run.

  • Routing the capacity refusal back through the rollback (dropping the
    capacity_refused branch): 3 kernel-mode arena tests fail, and both
    program-mode tests still pass.
  • Removing the retained-buffer guard: 2 tests fail, identically on a2a3 and
    a5
    ; the rest pass.
  • Dropping the addr != nullptr clause: exactly
    KernelModeAllocatesOnceThenRefusesToGrow fails
    , on both architectures — the
    first-bind regression has a precise barrier.
  • Removing the rollback entirely: exactly the two allocation-failure tests
    fail
    ; the nine capacity tests still pass, so the two failure classes are
    pinned independently of each other.

How the findings above were found

The commit was put through a self-review before this update: eight independent
reviewers over separate dimensions of the diff, then three adversarial refutation
lenses per finding. It raised 22 findings and the skeptic panel refuted all 22 —
a verdict I do not report as a clean bill of health, because the refutation lens
set included reachability, and since nothing latches KERNEL yet, every
kernel-mode finding is trivially unreachable today. The panel's value was in
surfacing candidates, not in adjudicating them.

Four of those candidates were textually verifiable and are fixed here regardless
of the vote: the first-allocation refusal, the is_kernel_mode comment claiming
table-wide address stability that acquire_graph_definition_block contradicts,
the stale contracts listed earlier, and the two test defects above — the
"smaller" run that packed to exactly the retained capacity, and the missing
rollback coverage.

🤖 Generated with Claude Code

Kernel mode is simpler's second execution identity: instead of owning
the device, a context borrows the caller's already-current device and
stream to enqueue one bounded asynchronous operator per launch, so a
PyPTO program is capturable by ACLGraph as an ordinary node. This
change freezes the public surface that identity hangs off and gives the
context a write-once identity the guards can key on. It creates no
resources. The program path gains one call — simpler_init latches
PROGRAM — and no behavior: latching a fresh context always succeeds, is
idempotent, and nothing on that path reads the latch.

- runtime_c_api.h declares the lifecycle entries
  simpler_kernel_mode_{supported,init,prepare_callable,launch} and adds
  the host-band code PTO_RUNTIME_ERR_INVALID_STATE. The existing
  finalize_device stays the fifth lifecycle entry, and a kernel context
  now reaches it. Kernel-mode capacity is
  a mode invariant rather than a gated state: config is context-static,
  so each pooled arena region is committed at most once, and
  setup_static_arena reports a grow or release request on a committed
  region under kernel mode as an internal invariant break; capacity
  intent travels in CallConfig.runtime_env like everywhere else.
- Execution identity is a write-once property of the context rather
  than a state that evolves. ExecutionModeLatch (platform/include/host/
  execution_mode_latch.h) replaces the four-state claim: the first init
  entry to run latches the mode, and it never changes — not on finalize,
  not on error. simpler_init latches PROGRAM before touching any
  process or runner state, so the program/kernel mutual exclusion is
  enforced on every program init instead of resting on a separate
  declaration call. There is no unlatch, which the latch documents as a
  consequence: a handle from a failed kernel init can never be recycled
  into a program context. SimplerExecutionMode now has one definition
  (task_interface/execution_mode.h) that both the wire header and the
  latch consume, so the host-side identity and the value that travels to
  the AICPU can no longer disagree.
- device_id_ records which device a context is on, not a claim on it —
  ownership is what the latch carries. attach_current_thread splits
  accordingly: bind_current_thread does the per-thread rtSetDevice and
  nothing else; attach_current_thread is the program-mode adopt
  (bind plus the one-shot op-execute watchdog and identity write) and
  refuses on a kernel latch; adopt_borrowed_device records the device a
  kernel context runs on without binding the thread and without
  configure_aicore_op_timeout, whose aclrtSetOpExecuteTimeOutV2 would
  rewrite the watchdog for every other user of a borrowed card. It does
  resolve the timeout config, because the stream and scheduler timeouts
  derived from it are read on both identities. DeviceRunner::finalize()
  is the one caller that runs under both identities and skips the bind
  on a kernel latch, so the kernel close path reaches its no-reset
  branch instead of being turned away by a device bind it never needed.
- ensure_acl_ready(), force_reset_device(), and finalize()'s rt-layer
  device reset refuse on a kernel-mode context (a2a3 + a5): the ACL
  lifecycle belongs to the caller, and every call site of the five ACL
  lifecycle APIs falls into three enumerable classes (below the
  ensure_acl_ready guard, inside force_reset_device behind its own
  guard, or gated on acl_ready_ which only the guarded path sets), with
  finalize's rt-layer reset intercepted by its own kernel-mode branch —
  so poison recovery can never reset the device out from under the
  host process.
- kernel_invocation_header.h pins the envelope every kernel launch
  ships to the AICPU (mode / callable / generation / payload length /
  int32_t arg counts). Both sides of the wire come from one
  build_runtimes.py build, so the struct carries no version or size
  negotiation and the POD/standard-layout guards are its only
  compile-time checks. generation is the occupancy counter of the
  residency slot callable_id resolves to - a property of the slot, not
  of the callable in it, so a generation carried by the callable could
  not detect slot reuse - with zero reserved for "not recorded".
  ChipCallable's sig_count includes the scalar entries and its
  scalar_count reads 0 both for a scalar-free orchestration and for an
  artifact built before the field existed, so a consumer derives the
  effective scalar count - the field when nonzero, otherwise the
  signature's SCALAR entries, the split count_callable_tensor_args
  already computes - and compares tensor_count against sig_count minus
  it. Subtracting the field directly would count an unrecorded
  callable's scalars as tensors.
- Kernel-entry argument validation is shared by all eight host-runtime
  components through kernel_entry_validation.h (one copy of the
  null/range/image-size/alignment checks; a binary pointer and its size
  must be present or absent together, and a callable image must be
  aligned for ChipCallable so its CALLABLE_CHILD_ALIGN-relative storage_
  lands aligned too), so a stub and a real implementation accept and
  reject exactly the same arguments.
- KernelExecutionState and ExecutionModeClaimState carry the kernel
  context phase machine (New/Collecting/ReadyEnqueued/Poisoned/
  Closing/Closed with sticky, retriable Closing and separate
  runtime-error and teardown-error slots) and the two restricted
  operation vocabularies; synchronize, allocation, capture queries,
  and model attachment stay unrepresentable in those tables, and a
  launch implementation is obligated to route through them. Every
  kernel-mode guard reads the identity through ExecutionModeLatch::
  is_kernel() rather than comparing an enumerator at the call site, so
  the test lives in one place instead of eight.
- ChipWorker dlsyms the four new symbols from every runtime, so a
  component missing one fails at load, and clears them alongside the
  other resolved pointers on all three teardown paths so none is left
  dangling into the library DlHandleGuard dlcloses.
  test_host_runtime_abi.py
  asserts the export across all eight components, and table-driven UTs
  cover the phase machine (including failed-rollback landing in
  Closing with the create error reported and the cleanup error
  latched), the shared argument validation, and the wire layout.
- ChipCallable additionally records scalar_count as a cached
  derivation of the signature's SCALAR entries: make_callable rejects
  a nonzero count that disagrees with the signature, while 0 also
  means "not recorded" (legacy blobs read 0). The field occupies four
  bytes of historical header tail padding, so every historical offset,
  sizeof, and the kernel-cache ABI token are unchanged;
  ChipCallable.build gains a trailing scalar_count=0 keyword and a
  read-only property.

Two facts a reader should not have to re-derive. The latch refusal returns
PTO_RUNTIME_ERR_INVALID_STATE (-1003) rather than PTO_RUNTIME_ERR_INTERNAL
(-1000) on purpose: conftest.py scrapes "simpler_init failed with code <N>"
and treats -1000 as a poisoned card, so an identity conflict must not look
like one. And kernel_execution_state.cpp stays compiled into all four host
runtimes even though grepping KernelExecutionState now finds only its own
header and .cpp — it is the persistent-state change's foundation, not an
orphaned translation unit.

Every kernel-mode branch this adds is provably dead in this commit: no
production site latches KERNEL (`git grep 'latch(SIMPLER_MODE_KERNEL)' src`
is empty) because both simpler_kernel_mode_init stubs return before any latch
call, so is_kernel() is false on every context and the program path takes the
same branch it took before.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 11, 2026

Copy link
Copy Markdown

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@sunkaixuan2018 sunkaixuan2018 changed the title Fix: kernel-mode capacity refusals no longer free what they protect (5K3) Fix: kernel-mode capacity refusals no longer free what they protect (⑤K3) Sep 11, 2026
@sunkaixuan2018
sunkaixuan2018 force-pushed the skx/kernel-capacity-freeze-k3 branch from eb69687 to 4297a35 Compare September 11, 2026 01:52
A kernel-mode context's device buffers are sized once and keep their
addresses, because a captured graph replays the addresses of the run it
captured; a re-base there produces silently wrong data rather than a
failure. Two refusals that enforce this were themselves destructive.

setup_static_arena's guard returned an error the caller collapsed into a
rollback that released all three regions, so a fired guard dropped the base
addresses it had just declined to move. The bank's commit rule now lives in
host/static_arena_bank.h, shared by the onboard and simulation runners
rather than duplicated in both, and it separates the two failures: an
allocation failure still rolls the whole bank back, while a capacity refusal
leaves every region committed and every cached size intact.

RetainedTempBump::begin freed the retained buffer before asking for a larger
one, so a refusal could only land after the address was already gone. It now
refuses ahead of the free. The refusal is scoped to a slot that already
holds a buffer: an empty slot has no address for a captured graph to hold,
so a context's first allocation is not a re-base and is taken normally,
which is what lets a kernel context reach its frozen size at all. Nothing in
the tree restricts a kernel context's tensors to device memory, so that
first allocation is reachable.

The context's identity reaches the runtime through a new HostApiOps entry,
is_kernel_mode. It is a query of the existing latch, not a new gate, and a
table that does not supply it reports program mode.

Program mode keeps growth, release and rollback unchanged. Both arena call
sites keep their prebuilt-arena cache invalidation, now keyed on whether the
commit moved a base.

The contracts that described the behavior these guards make conditional move
with them: the setup_static_arena and retained-temp-buffer entries in
host_api.h, DeviceRunnerBase::setup_static_arena's rollback promise and its
return code, the kernel-mode capacity paragraph in runtime_c_api.h, both
runners' latch comments, task-flow.md, and RUNTIME_LOGIC.md in both
architecture trees.

tests/ut/cpp/common/test_static_arena_bank.cpp covers the bank rule with no
device: base and capacity constant across repeated calls, zero allocator
calls, exact capacity accepted, one byte over refused with
PTO_RUNTIME_ERR_INTERNAL, and the held layout still served afterwards. The
TRB suite gains the same coverage on the retained temporary buffer for both
architectures, including a context that is in kernel mode from its first
bind, which is the only shape the write-once latch permits. Both suites pin
the other exit too: an allocation failure rolls the whole bank back, in
kernel mode as in program mode, since a setup that never completed published
no address to protect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@ChaoWao

ChaoWao commented Sep 13, 2026

Copy link
Copy Markdown
Collaborator

上板实测,正好落在这个 PR 的前提上

#2064 已合(1d1ddc81),这条是把我在 K1 上提过的东西里只与本 PR 相关的部分挪过来,外加三条新的硬件实测。

环境:a2a3 onboard、CANN 9.0.0、torch_npu 2.7.1。代码与全部结果在 docs-kernel-mode/probes/(每条附复跑命令与观测边界)。


1. 你的核心前提"silently wrong data on replay, not a crash"——CANN 侧实测证实,而且比想象的更彻底

我用 raw aclmdlRICapture* 直驱、每个操作一次独立 capture、用 aclmdlRICaptureGetInfo 区分"当场拒绝/被捕获/静默污染"。结果里与本 PR 直接相关的一行:

操作(捕获期间) op rc capture status CaptureEnd 出图
aclrtMalloc 0 ACTIVE 0
aclrtFree 0 ACTIVE 0

GLOBAL 与 RELAXED 两种捕获模式一致。 也就是说:一次 re-base 发生在被捕获的 launch 里,CANN 不报错、不把 capture 标脏、照常出图

没有任何运行期兜底。这个 PR 的 guard 就是唯一挡在"re-base"和"replay 读错地址"之间的东西。 我认为这句可以直接写进 PR 描述——它把这条从"防御性编程"抬成"唯一防线"。

2. 对照:CANN 确实会拦的那些,恰好不含分配

同一批探针里被当场拒绝的操作,说明 CANN 的捕获校验并非形同虚设:

操作 错误码
wait 一个在捕获开始前 record 的 event 107024
fork 到侧流但不 join 回来(CaptureEnd 时) 107025
aclrtSynchronizeStream(捕获流)/SynchronizeDeviceStreamQuery(捕获流) 107027
aclrtSynchronizeEvent / aclrtQueryEventStatus 107028
aclrtMemcpy / aclrtMemset(同步) 107030

所以"地址与容量"正是那一类零执法的不变式,而同步、事件误用、悬空 fork 都有码可报。这个对照我觉得比单说"没人管"更有说服力:不是 CANN 疏漏了整块,是分配这一类它根本不视为捕获相关操作

3. "an ACLGraph replays the addresses that were captured"——两侧都测过了

正面:两张图共享同一对 device slot、各写自己的 pattern,交替 replay A B A B A A B 7/7背靠背不同步 replay(AB/BA/AA/BB)4/4,每次 slot 恰好翻成该图的 pattern。replay 确实照着捕获时的地址写。

反面(更贴近你说的失败形态):把工作挂在一条没有进入本次捕获的流上——不进图、不报错、replay 静默地什么都不做,只有数值不对。零 ACL 错误。

这两条合起来就是你 PR 里那句话的实测版本:错的地址不会让你崩,只会让你算错。


4. 一个问题:第一次分配允许发生在 capture 之内吗?

你把 addr != nullptr 这半说清楚了,我同意窄 guard 是对的。顺着问一句:

按 §1,捕获期间 aclrtMalloc放行的,而且分配不是流上的操作、不进图。所以如果某个 kernel context 的第一次 bind 恰好发生在一次被捕获的 launch 里,这次分配会真实发生、地址被后续录进图、且因为再也不 re-base 而保持稳定——结果是对的,但"capture 期做了一次真实分配"这件事本身值得明说是允许还是应当被 prepare 期兜住。

如果设计上第一次分配必须落在 prepare(capture 之外),那它就不只是"guard 放行",而是⑫ST 该有一条负例:capture 内出现首次分配 → 必须被检出。我没有主张改代码,只是这条今天没有任何东西挡。

5. rebase 之后请重锚

本 PR 基于 86dd62b4。此后除 K1 合入外,#2163 / #2200 / #2201 / #2204 四个 DFX PR 大改了 device_runner_base.cpp 与两个 arch 的 device_runner.cpp,我量过的位移:

b28c4c4d 现在
start_shared_collectors_for_run 定义 1888 1930(且改收 (const DfxRunConfig&, uint32_t pipeline_slot)
teardown_shared_collectors_after_run 1924 1975
ensure_device_wall_buffer 1654 1696
arm_device_wall_buffer 1814 1722
DFX 位图进 kernel_args device_runner.cpp:292-302 :1147-1150

allow_prepared_successor 里的 && !config->diagnostics_any() 也已删除(c_api_shared.cpp:800)。文件动得比较厉害,rebase 后值得整体过一遍而不只看冲突标记。


不属于本 PR 的:我在 K1 上提的"潜伏缺陷修在哪"那条规则,对应的是 ⑩a / #2185(它才是让 kernel context 真正可构造的那个),不在这里。本 PR 是已合入 guard 的正确性修复,位置我认为是对的。

边界:以上全部限 a2a3,a5 一个没跑——按非目标 #12 不得由 a2a3 推断 a5。

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants