<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>故障排查 on 闫工的技术博客</title>
    <link>https://www.cciepro.cn/tags/%E6%95%85%E9%9A%9C%E6%8E%92%E6%9F%A5/</link>
    <description>Recent content in 故障排查 on 闫工的技术博客</description>
    <generator>Hugo -- gohugo.io</generator>
    <language>zh-cn</language>
    <copyright>© 2026 闫工 · [宁ICP备2026002305号-3](https://beian.miit.gov.cn/)</copyright>
    <lastBuildDate>Wed, 26 Aug 2026 07:00:00 +0800</lastBuildDate><atom:link href="https://www.cciepro.cn/tags/%E6%95%85%E9%9A%9C%E6%8E%92%E6%9F%A5/index.xml" rel="self" type="application/rss+xml" />
    
    <item>
      <title>8卡变7卡、93°C还在Throttling：卡没坏，是热到&#39;自我降速&#39;了</title>
      <link>https://www.cciepro.cn/posts/gpu-throttling-93c/</link>
      <pubDate>Wed, 26 Aug 2026 07:00:00 +0800</pubDate>
      
      <guid>https://www.cciepro.cn/posts/gpu-throttling-93c/</guid>
      <description>凌晨训练任务挂了，8 卡只剩 7 卡，温度 93°C 还带 Throttling——卡没坏，是太热了。从确认告警、风道检查、清灰换硅脂到 vBIOS 更新，完整排查流程，满载压回 80°C。</description>
      
    </item>
    
    <item>
      <title>GPU全在、驱动正常，训练就是卡死——NVLink悄悄坏了一半</title>
      <link>https://www.cciepro.cn/posts/nvlink-half-broken-training-stuck/</link>
      <pubDate>Sun, 23 Aug 2026 10:00:03 +0800</pubDate>
      
      <guid>https://www.cciepro.cn/posts/nvlink-half-broken-training-stuck/</guid>
      <description>8 卡 A800 训练卡死，显卡全在、驱动正常，可训练就是跑不动。查下来是 GPU 之间的 NVLink 高速通道废了 5 条。改一行配置让服务&amp;quot;带伤上岗&amp;quot;，集群满血复活。</description>
      
    </item>
    
    <item>
      <title>服务器红灯告警说硬盘坏了，我没换——结果真救活了</title>
      <link>https://www.cciepro.cn/posts/raid-false-alarm-hdd/</link>
      <pubDate>Fri, 31 Jul 2026 07:30:00 +0800</pubDate>
      
      <guid>https://www.cciepro.cn/posts/raid-false-alarm-hdd/</guid>
      <description>服务器报警说硬盘坏了，监控群里喊换盘。我没换——因为那多半是 RAID 卡误报的「疑似坏盘」标记，盘根本没坏。用 StorCLI 5 分钟平反：先看全局、再查 SMART 和健康日志，逻辑问题一键重建，真有物理坏道才换。别让一块好盘白白被换掉。</description>
      
    </item>
    
    <item>
      <title>凌晨3点的GPU服务器：一次BMC故障排查全记录</title>
      <link>https://www.cciepro.cn/posts/gpu-bmc-troubleshooting-3am/</link>
      <pubDate>Sun, 26 Jul 2026 07:30:00 +0800</pubDate>
      
      <guid>https://www.cciepro.cn/posts/gpu-bmc-troubleshooting-3am/</guid>
      <description>凌晨3点，集群一台8卡GPU节点硬掉线。SSH连不上、ping不通，但BMC还活着。翻SEL发现供电在抖，锁定了坏掉的PSU3，顺带挖出NVIDIA驱动和内核的兼容性雷。不卖工具，卖教训。</description>
      
    </item>
    
  </channel>
</rss>
