Repository navigation
[Bug] 1.3.6服务运行突然configNode从Runing变成了disabled 没有原因日志 #18537
Description
Activity
可以发一下完整日志
configNode的日志:
2026-08-28 00:06:14,847 [ProcedureTimeoutExecutor] INFO o.a.i.c.p.PartitionTableAutoCleaner:66 - [PartitionTableCleaner] Periodically activate PartitionTableAutoCleaner, databaseTTL: {root.db_equipment_status=-1, root.db_location=-1, root.db_operation_log=-1, root.db_scene_linkage_log=-1, root.db_statistic_data=-1, root.db_system_log=-1, root.db_user_login_log=-1, root.ln_1=-1, root.ln_2=-1, root.ln_event_original=-1, root.operation_log=-1}
2026-08-28 02:06:14,849 [ProcedureTimeoutExecutor] INFO o.a.i.c.p.PartitionTableAutoCleaner:66 - [PartitionTableCleaner] Periodically activate PartitionTableAutoCleaner, databaseTTL: {root.db_equipment_status=-1, root.db_location=-1, root.db_operation_log=-1, root.db_scene_linkage_log=-1, root.db_statistic_data=-1, root.db_system_log=-1, root.db_user_login_log=-1, root.ln_1=-1, root.ln_2=-1, root.ln_event_original=-1, root.operation_log=-1}
2026-08-28 04:06:14,849 [ProcedureTimeoutExecutor] INFO o.a.i.c.p.PartitionTableAutoCleaner:66 - [PartitionTableCleaner] Periodically activate PartitionTableAutoCleaner, databaseTTL: {root.db_equipment_status=-1, root.db_location=-1, root.db_operation_log=-1, root.db_scene_linkage_log=-1, root.db_statistic_data=-1, root.db_system_log=-1, root.db_user_login_log=-1, root.ln_1=-1, root.ln_2=-1, root.ln_event_original=-1, root.operation_log=-1}
2026-08-28 05:46:46,254 [pool-12-IoTDB-Cluster-Event-Service-1] INFO o.a.i.c.m.l.s.EventService:202 - [RegionGroupStatistics] RegionGroupStatisticsMap:
2026-08-28 06:54:28,132 [pool-12-IoTDB-Cluster-Event-Service-1] INFO o.a.i.c.m.l.s.EventService:207 - [RegionGroupStatistics] RegionGroup TConsensusGroupId(type:SchemaRegion, id:9): Running -> DisableddataNode的日志:
[root@ecm-b397-0004 iotdb]# cat log_datanode_all.log |less
2026-08-28 05:26:00,893 [pool-36-IoTDB-Timed-Flush-Seq-Memtable-1] INFO o.a.i.d.s.d.DataRegion:1721 - Exceed sequence memtable flush interval, so flush working memtable of time partition 2956 in database root.ln_2[10]
2026-08-28 09:09:52,474 [JvmPauseMonitor0] WARN o.a.r.u.JvmPauseMonitor:168 - JvmPauseMonitor-1: Detected pause in JVM or host machine approximately 5.304s without any GCs.
2026-08-28 09:09:52,830 [pool-36-IoTDB-Timed-Flush-Seq-Memtable-1] INFO o.a.i.d.s.d.DataRegion:1598 - Async close tsfile: /iotdb/data/datanode/data/sequence/root.ln_2/10/2956/1787825242546-2-0-0.tsfile, file start time: 1787825242132, file end time: 1787865142357
2026-08-28 09:09:52,830 [pool-36-IoTDB-Timed-Flush-Seq-Memtable-1] INFO o.a.i.d.s.d.DataRegion:1721 - Exceed sequence memtable flush interval, so flush working memtable of time partition 2956 in database root.ln_1[12]
2026-08-28 09:09:52,831 [pool-36-IoTDB-Timed-Flush-Seq-Memtable-1] INFO o.a.i.d.s.d.DataRegion:1598 - Async close tsfile: /iotdb/data/datanode/data/sequence/root.ln_1/12/2956/1787825242582-2-0-0.tsfile, file start time: 1787825242132, file end time: 1787865142357
2026-08-28 09:09:52,924 [pool-58-IoTDB-ClientRPC-Processor-5$20260828_010952_00000_1] INFO o.a.i.c.c.ThriftClient:93 - Broken pipe error happened in sending RPC, we need to clear all previous cached connection, error msg is java.net.ConnectException: Connection refused
2026-08-28 09:09:52,925 [pool-58-IoTDB-ClientRPC-Processor-5$20260828_010952_00000_1] WARN o.a.i.d.p.c.ConfigNodeClient:283 - The current node leader may have been down TEndPoint(ip:127.0.0.1, port:10710), try next node
2026-08-28 09:09:52,925 [pool-58-IoTDB-ClientRPC-Processor-5$20260828_010952_00000_1] INFO o.a.i.c.c.ThriftClient:93 - Broken pipe error happened in sending RPC, we need to clear all previous cached connection, error msg is java.net.ConnectException: Connection refused
2026-08-28 09:09:52,925 [pool-58-IoTDB-ClientRPC-Processor-5$20260828_010952_00000_1] WARN o.a.i.d.p.c.ConfigNodeClient:308 - The current node may have been down TEndPoint(ip:127.0.0.1, port:10710),try next node
2026-08-28 09:09:52,933 [pool-58-IoTDB-ClientRPC-Processor-2$20260828_010952_00001_1] INFO o.a.i.c.c.ThriftClient:93 - Broken pipe error happened in sending RPC, we need to clear all previous cached connection, error msg is java.net.ConnectException: Connection refused
2026-08-28 09:09:52,933 [pool-58-IoTDB-ClientRPC-Processor-2$20260828_010952_00001_1] WARN o.a.i.d.p.c.ConfigNodeClient:283 - The current node leader may have been down TEndPoint(ip:127.0.0.1, port:10710), try next node
2026-08-28 09:09:52,933 [pool-58-IoTDB-ClientRPC-Processor-2$20260828_010952_00001_1] INFO o.a.i.c.c.ThriftClient:93 - Broken pipe error happened in sending RPC, we need to clear all previous cached connection, error msg is java.net.ConnectException: Connection refused
2026-08-28 09:09:52,933 [pool-58-IoTDB-ClientRPC-Processor-2$20260828_010952_00001_1] WARN o.a.i.d.p.c.ConfigNodeClient:308 - The current node may have been down TEndPoint(ip:127.0.0.1, port:10710),try next node
2026-08-28 09:09:53,099 [pool-14-IoTDB-Flush-2] INFO o.a.i.d.s.d.m.TsFileProcessor:1561 - The compression ratio of tsfile /iotdb/data/datanode/data/sequence/root.ln_1/12/2956/1787825242582-2-0-0.tsfile is 12.93, totalMemTableSize: 13320, the file size: 1030
2026-08-28 09:09:53,151 [pool-14-IoTDB-Flush-1] INFO o.a.i.d.s.d.m.TsFileProcessor:1561 - The compression ratio of tsfile /iotdb/data/datanode/data/sequence/root.ln_2/10/2956/1787825242546-2-0-0.tsfile is 12.93, totalMemTableSize: 13320, the file size: 1030
2026-08-28 09:09:53,926 [pool-58-IoTDB-ClientRPC-Processor-5$20260828_010952_00000_1] INFO o.a.i.c.c.ThriftClient:93 - Broken pipe error happened in sending RPC, we need to clear all previous cached connection, error msg is java.net.ConnectException: Connection refused
2026-08-28 09:09:53,926 [pool-58-IoTDB-ClientRPC-Processor-5$20260828_010952_00000_1] WARN o.a.i.d.p.c.ConfigNodeClient:308 - The current node may have been down TEndPoint(ip:127.0.0.1, port:10710),try next node
2026-08-28 09:09:53,927 [pool-50-IoTDB-MPP-Coordinator-Executor-20] WARN o.a.i.d.q.p.e.c.ConfigExecution:123 - Failures happened during running ConfigExecution.
org.apache.iotdb.commons.client.exception.ClientManagerException: net.sf.cglib.core.CodeGenerationException: org.apache.thrift.TException-->Fail to connect to any config node. Please check status of ConfigNodes or logs of connected DataNode
at org.apache.iotdb.commons.client.ClientManager.borrowClient(ClientManager.java:57)
at org.apache.iotdb.db.queryengine.plan.execution.config.executor.ClusterConfigTaskExecutor.showDataNodes(ClusterConfigTaskExecutor.java:1443)
at org.apache.iotdb.db.queryengine.plan.execution.config.metadata.ShowDataNodesTask.execute(ShowDataNodesTask.java:56)
at org.apache.iotdb.db.queryengine.plan.execution.config.ConfigExecution.start(ConfigExecution.java:98)
at org.apache.iotdb.db.queryengine.plan.Coordinator.execution(Coordinator.java:151)
at org.apache.iotdb.db.queryengine.plan.Coordinator.executeForTreeModel(Coordinator.java:187)
at org.apache.iotdb.db.protocol.thrift.impl.ClientRPCServiceImpl.executeStatementInternal(ClientRPCServiceImpl.java:317)
at org.apache.iotdb.db.protocol.thrift.impl.ClientRPCServiceImpl.executeStatementV2(ClientRPCServiceImpl.java:787)
at org.apache.iotdb.db.protocol.thrift.impl.ClientRPCServiceImpl.executeQueryStatementV2(ClientRPCServiceImpl.java:777)
Search before asking
Version
1.3.6
Describe the bug and provide the minimal reproduce step
1.3.6服务运行突然configNode从Runing变成了disabled 没有原因日志
What did you expect to see?
正常运行,或者能看到dataNode由Runing变成了disabled的原因
What did you see instead?
突然dataNode从Runing变成了disabled
Anything else?
No response
Are you willing to submit a PR?