Skip to content

Extend the node state machine and persist node statuses - #18801

Merged
CRZbulabula merged 5 commits into
apache:masterfrom
d-wang-commit:extend-and-persist-node-state-machine
Oct 9, 2026
Merged

CRZbulabula merged 5 commits into
apache:masterfrom
d-wang-commit:extend-and-persist-node-state-machine

Conversation

@d-wang-commit

Copy link
Copy Markdown
Contributor

Description

Durable shutdown and removal states

Distinguish reported shutdowns from heartbeat failures with Stopped for ConfigNodes, DataNodes and AINodes. Persist Stopped and Removing through ConfigNode consensus and snapshots so leader changes and restarts retain these states. Clear obsolete records on recovery or node deregistration.

Centralize status transitions: ordinary updates preserve Removing, heartbeat failures cannot replace Stopped with Unknown, and live heartbeats can revive a stopped node. Explicit management updates support removal rollback. Update routing, scheduling and removal checks to handle offline nodes consistently.

Use deterministic SET_STOPPED, SET_REMOVING and CLEAR commands. Calculate each transition once and avoid consensus writes when the durable record already matches. A failed write is returned to the caller while current in-memory statistics still advance; subsequent statistics updates reconcile the durable record. Shutdown acknowledgements, removal progress and leader readiness check the required persistence result.

ReadOnly reasons and node displays

Classify local ReadOnly reasons as DiskFull, UnrecoverableError, Manual or Stopping. For repeated local ReadOnly updates, apply Stopping > Manual > UnrecoverableError > DiskFull and preserve the first reason at equal priority. Automatic disk recovery only clears DiskFull; received heartbeats retain the remote node's chosen reason.

Expose status and reason separately through node RPCs and information_schema.nodes. Append an optional StatusReason column to SHOW CLUSTER, SHOW CLUSTER DETAILS, SHOW DATANODES and SHOW CONFIGNODES in both SQL dialects when a reason is present. Rows without a reason contain NULL.

ReadOnly and its reason remain transient and are rebuilt from live heartbeats after a leader change. A revived node can still require a consensus CLEAR to remove an old Stopped record.

Tests

  • Added or extended UTs for transitions, persistence failures and retries, shutdown reports, removal rollback, log/snapshot state handling, reason precedence and propagation, disk recovery, and SHOW layouts. Key suites include NodeStatusTest, LoadCachePersistedNodeStatusTest, NodeInfoTest, ConfigNodeShutdownHookTest, RemoveDataNodePersistenceTest, CommonConfigTest and the four Show*TaskTest classes. Added AINode shutdown tests in tests/test_shutdown.py.
  • Added or extended ITs for shutdown, leader failover, node removal, log/snapshot recovery, AINode persistence and reason display. IoTDBSetSystemStatusTableIT now checks all four SHOW commands through tree and table connections, including the layout after returning to Running.
  • Selected Java UTs passed with no failures or errors; four existing immutable-directory cases were skipped because they require manual setup. English and Chinese full-reactor compilation passed.
  • Four targeted IT methods passed: IoTDBClusterNodeShutdownHookIT#testNodeShutdownReporter, IoTDBNodeStatusFailoverIT#testRealRemovalPersistsIntentAndDeletesRegistration, IoTDBNodeStatusPersistenceIT#testReadOnlyReasonRebuiltAfterLeaderFailureAndClearedAfterCrash, and IoTDBSetSystemStatusTableIT#setSystemStatus.
  • These results cover the selected suites, not every added test or the entire regression suite. Windows local-stop behavior is accounted for in the shutdown IT.

Side effects and risks

  • Clients must handle the new Stopped status and consume reasons separately from status strings. SHOW preserves existing column positions, but its column count can change when reasons appear. The fixed status_reason column in information_schema.nodes shifts subsequent positions for SELECT * consumers.
  • Older snapshots without persisted node statuses remain readable. The new consensus plan and AINode shutdown RPC require compatible peers; mixed-version ConfigNode operation has not been validated.
  • Persistent status changes add consensus writes and a small status map/snapshot section. Slow writes hold the affected node's sample lock and can delay its updates and the periodic statistics pass. On write failure, displayed in-memory state can temporarily differ from the durable record; callers receive the failure and later refreshes retry reconciliation.

This PR has:

  • been self-reviewed.
  • added comments explaining the "why" and the intent of the code wherever would not be obvious
    for an unfamiliar reader.
  • added integration tests.
  • been tested in a test IoTDB cluster.

Key changed/added classes (or packages if there are too many classes) in this PR
  • NodeStatus and NodeStatistics: shared transition rules, status predicates and reason handling.
  • LoadManager, LoadCache and node heartbeat caches: status updates, persistence callbacks, recovery and statistics publication.
  • UpdateNodeStatusPlan, ConfigPlanExecutor and NodeInfo: consensus commands and durable status storage, snapshots and cleanup.
  • ConfigManager, shutdown hooks and AINode RPC client/handler: shutdown reporting and acknowledgements.
  • Removal procedures, region allocation and routing code: consistent offline/removing decisions and retryable status updates.
  • CommonConfig, DataNodeInternalRPCServiceImpl, disk strategies and storage error handlers: local ReadOnly reasons and disk recovery.
  • NodeManager, Show*Task, InformationSchemaContentSupplierFactory and Thrift definitions: separate status/reason values in node displays.

Report graceful shutdown as Stopped, apply shared node status transition rules, and persist Stopped/Removing statuses through ConfigNode consensus for recovery after leadership changes.
Cover shutdown reporting, status transitions, scheduling and removal behavior, and node status persistence and recovery with unit and integration tests.
Classify local ReadOnly reasons and propagate them through heartbeats and node display APIs.
Keep reasons separate from status and append StatusReason to SHOW results when present.
Cover reason precedence, explicit updates, removal rollback, SHOW layouts and table node metadata.
Verify reason recovery after leader changes and cleanup after node failure.
@d-wang-commit
d-wang-commit force-pushed the extend-and-persist-node-state-machine branch from fadbba6 to a7bb3cd Compare October 9, 2026 03:18

@CRZbulabula CRZbulabula left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Comment thread iotdb-core/ainode/iotdb/ainode/core/rpc/client.py Outdated
@CRZbulabula
CRZbulabula merged commit 94f99d6 into apache:master Oct 9, 2026
38 of 39 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants