Repository navigation
Extend the node state machine and persist node statuses - #18801
Merged
CRZbulabula merged 5 commits intoOct 9, 2026
Merged
CRZbulabula merged 5 commits into
CRZbulabula merged 5 commits into
Conversation
Report graceful shutdown as Stopped, apply shared node status transition rules, and persist Stopped/Removing statuses through ConfigNode consensus for recovery after leadership changes.
Cover shutdown reporting, status transitions, scheduling and removal behavior, and node status persistence and recovery with unit and integration tests.
Classify local ReadOnly reasons and propagate them through heartbeats and node display APIs. Keep reasons separate from status and append StatusReason to SHOW results when present.
Cover reason precedence, explicit updates, removal rollback, SHOW layouts and table node metadata. Verify reason recovery after leader changes and cleanup after node failure.
d-wang-commit
force-pushed
the
extend-and-persist-node-state-machine
branch
from
October 9, 2026 03:18
fadbba6 to
a7bb3cd
Compare
CRZbulabula
approved these changes
Oct 9, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Durable shutdown and removal states
Distinguish reported shutdowns from heartbeat failures with
Stoppedfor ConfigNodes, DataNodes and AINodes. PersistStoppedandRemovingthrough ConfigNode consensus and snapshots so leader changes and restarts retain these states. Clear obsolete records on recovery or node deregistration.Centralize status transitions: ordinary updates preserve
Removing, heartbeat failures cannot replaceStoppedwithUnknown, and live heartbeats can revive a stopped node. Explicit management updates support removal rollback. Update routing, scheduling and removal checks to handle offline nodes consistently.Use deterministic
SET_STOPPED,SET_REMOVINGandCLEARcommands. Calculate each transition once and avoid consensus writes when the durable record already matches. A failed write is returned to the caller while current in-memory statistics still advance; subsequent statistics updates reconcile the durable record. Shutdown acknowledgements, removal progress and leader readiness check the required persistence result.ReadOnly reasons and node displays
Classify local ReadOnly reasons as
DiskFull,UnrecoverableError,ManualorStopping. For repeated local ReadOnly updates, applyStopping > Manual > UnrecoverableError > DiskFulland preserve the first reason at equal priority. Automatic disk recovery only clearsDiskFull; received heartbeats retain the remote node's chosen reason.Expose status and reason separately through node RPCs and
information_schema.nodes. Append an optionalStatusReasoncolumn toSHOW CLUSTER,SHOW CLUSTER DETAILS,SHOW DATANODESandSHOW CONFIGNODESin both SQL dialects when a reason is present. Rows without a reason contain NULL.ReadOnly and its reason remain transient and are rebuilt from live heartbeats after a leader change. A revived node can still require a consensus
CLEARto remove an old Stopped record.Tests
NodeStatusTest,LoadCachePersistedNodeStatusTest,NodeInfoTest,ConfigNodeShutdownHookTest,RemoveDataNodePersistenceTest,CommonConfigTestand the fourShow*TaskTestclasses. Added AINode shutdown tests intests/test_shutdown.py.IoTDBSetSystemStatusTableITnow checks all four SHOW commands through tree and table connections, including the layout after returning to Running.IoTDBClusterNodeShutdownHookIT#testNodeShutdownReporter,IoTDBNodeStatusFailoverIT#testRealRemovalPersistsIntentAndDeletesRegistration,IoTDBNodeStatusPersistenceIT#testReadOnlyReasonRebuiltAfterLeaderFailureAndClearedAfterCrash, andIoTDBSetSystemStatusTableIT#setSystemStatus.Side effects and risks
Stoppedstatus and consume reasons separately from status strings. SHOW preserves existing column positions, but its column count can change when reasons appear. The fixedstatus_reasoncolumn ininformation_schema.nodesshifts subsequent positions forSELECT *consumers.This PR has:
for an unfamiliar reader.
Key changed/added classes (or packages if there are too many classes) in this PR
NodeStatusandNodeStatistics: shared transition rules, status predicates and reason handling.LoadManager,LoadCacheand node heartbeat caches: status updates, persistence callbacks, recovery and statistics publication.UpdateNodeStatusPlan,ConfigPlanExecutorandNodeInfo: consensus commands and durable status storage, snapshots and cleanup.ConfigManager, shutdown hooks and AINode RPC client/handler: shutdown reporting and acknowledgements.CommonConfig,DataNodeInternalRPCServiceImpl, disk strategies and storage error handlers: local ReadOnly reasons and disk recovery.NodeManager,Show*Task,InformationSchemaContentSupplierFactoryand Thrift definitions: separate status/reason values in node displays.