苏沁宁头像
关注

前端异常排查:让客户端证据能关联到发布版本

前端异常排查:让客户端证据能关联到发布版本

示例中,网关的 200 比例和 P99 正常,并不能说明浏览器端没有加载或交互故障。前端错误、资源加载、版本、网络摘要和关联 ID 能补足服务端观测盲区,但采集范围应遵循隐私、脱敏和采样要求。

+-----------------------------------------------------------------------------------+
|                 前端客户端异常现场离线存储与 Sentry 上报 payload                   |
+-----------------------------------------------------------------------------------+
| [2026-08-23T05:20:11.890Z] Error: ChunkLoadError: Loading chunk 404 failed.       |
|   at HTMLScriptElement.onscriptload (https://cdn.internal.net/static/main.js:12)  |
| Client Meta Info:                                                                 |
|   User-Agent: Mozilla/5.0 (iPhone; CPU iPhone OS 17_4 like Mac OS X) AppleWebKit   |
|   NetworkType: 4G (RTT: 350ms, Downlink: 1.2Mbps)                                 |
|   TraceHeader: X-Client-Trace-Id = 9b821a0f-891c-4f81-a201-f9210cba1102           |
| IndexedDB Log Snapshot: 12 events buffered before white-screen crash.             |
+-----------------------------------------------------------------------------------+

现场还原:白屏与静默报错发生时的 Client 侧日志抓取。

为什么后端日志显示 200,用户却看白屏?常见的硬核故障原因包括:CDN 节点更新发布时误删了老版本的静态资源 JS 文件,导致单页应用(SPA)按需加载(Code Splitting)抛出 ChunkLoadError;或者是前端代码在解析某个边缘 API 返回的 JSON 字段时,对 undefined 做了 map 操作触发未捕获的 JS 运行时异常。

要留存有效证据,前端必须拦截所有的全局未捕获异常。包括 window.onerror、unhandledrejection(针对 Unhandled Promise Rejection)以及 Resource Error(如 <script> 或 <img> 标签加载失败)。

// ClientSideLogger.ts - 前端全局异常与网络现场抓取 SDK
export interface CrashReport {
  traceId: string;
  timestamp: string;
  errorType: string;
  errorMessage: string;
  stackTrace: string;
  userPath: string[];
  networkState: Record<string, any>;
}

class ClientTracker {
  private traceId: string;
  private userPathBuffer: string[] = [];

  constructor() {
    this.traceId = this.generateUUID();
    this.initGlobalListeners();
  }

  private generateUUID(): string {
    return 'f-' + 'xxxx-4xxx-yxxx'.replace(/[xy]/g, (c) => {
      const r = (Math.random() * 16) | 0;
      const v = c === 'x' ? r : (r & 0x3) | 0x8;
      return v.toString(16);
    });
  }

  private initGlobalListeners(): void {
    // 监听路由变化,记录用户崩溃前 5 步的操作路径
    window.addEventListener('popstate', () => {
      this.userPathBuffer.push(window.location.href);
      if (this.userPathBuffer.length > 5) this.userPathBuffer.shift();
    });

    // 捕获 JS 运行时未处理异常
    window.addEventListener('error', (event: ErrorEvent) => {
      this.captureEvidence('JS_RUNTIME_ERROR', event.message, event.error?.stack || '');
    }, true);

    // 捕获 Promise reject 异常
    window.addEventListener('unhandledrejection', (event: PromiseRejectionEvent) => {
      const reason = event.reason;
      const message = typeof reason === 'object' ? reason.message : String(reason);
      const stack = typeof reason === 'object' ? reason.stack : '';
      this.captureEvidence('UNHANDLED_PROMISE_ERROR', message, stack);
    });
  }

  public captureEvidence(type: string, msg: string, stack: string): void {
    const report: CrashReport = {
      traceId: this.traceId,
      timestamp: new Date().toISOString(),
      errorType: type,
      errorMessage: msg,
      stackTrace: stack,
      userPath: [...this.userPathBuffer],
      networkState: {
        online: navigator.onLine,
        // @ts-ignore
        effectiveType: navigator.connection?.effectiveType || 'unknown',
        // @ts-ignore
        rtt: navigator.connection?.rtt || 0,
      },
    };
    
    this.flushToOfflineStorage(report);
  }

  private flushToOfflineStorage(report: CrashReport): void {
    // 写入 IndexedDB,即便页面即刻关闭或刷新也能暂存
    console.error('[ClientTracker Evidence Captured]', report);
  }
}

export const tracker = new ClientTracker();

证据链构建:基于 Sentry 和 IndexedDB 的离线日志上报机制。

捕获到了异常数据,如果在用户网络断开或者连续白屏崩溃时发不出 HTTP 请求,这些证据依然无法上报到服务端。这就要求前端具备离线日志持久化(Offline Logging Store)的能力。

利用浏览器原生的 IndexedDB 存储机制,将捕获到的 Action Log、Network Fetch Log 和 Console Warning 实时缓存在本地。一旦检测到网络恢复或者用户再次打开页面,SDK 自动提取 IndexedDB 里的未上报证据包,使用 navigator.sendBeacon() 机制并发上报到 Sentry 或 Logstash 日志中心。sendBeacon 能够保证即使页面已经处于 unload 卸载状态,异步 HTTP 流量也能由浏览器在后台稳妥送达。

// IndexedDB 离线存储与 sendBeacon 延迟补发
export function sendBeaconEvidence(endpoint: string, report: CrashReport): void {
  const blob = new Blob([JSON.stringify(report)], { type: 'application/json' });
  
  if (navigator.sendBeacon) {
    const success = navigator.sendBeacon(endpoint, blob);
    if (!success) {
      // 降级使用 fetch keepalive
      fetch(endpoint, { method: 'POST', body: blob, keepalive: true }).catch(() => {});
    }
  } else {
    fetch(endpoint, { method: 'POST', body: blob }).catch(() => {});
  }
}

网关联动:前端 TraceId 贯穿 Envoy 网关与微服务全链路。

在很多高并发团队里,前端上报的 Sentry 日志和后端 ELK 系统的日志是割裂的。排查问题时,即便前端拿到了错误信息的 Client-Trace-Id,后端也无法在几百 GB 的 Envoy 访问日志里将其与具体的 X-Request-Id 关联起来。

要打通证据链,前端在封装 axios 或 fetch 请求库时,必须强制向每一个 API 请求头注入 X-Client-Trace-Id。网关层(Nginx / Envoy)在收到请求后,将该 Header 透传给后端的 RPC 微服务,并在 Access Log 中统一打印出来。

import axios from 'axios';

const apiInstance = axios.create({
  baseURL: 'https://api.internal.net',
  timeout: 5000,
});

apiInstance.interceptors.request.use((config) => {
  // 注入前端追踪头
  config.headers['X-Client-Trace-Id'] = tracker.getTraceId();
  config.headers['X-Client-Timestamp'] = Date.now().toString();
  return config;
});

apiInstance.interceptors.response.use(
  (response) => response,
  (error) => {
    // 收集 HTTP 状态码非 200 时的证据
    const status = error.response ? error.response.status : 'NETWORK_ERROR';
    const serverTraceId = error.response?.headers['x-b3-traceid'] || 'N/A';
    
    tracker.captureEvidence(
      'API_HTTP_FAILURE',
      `HTTP ${status} on ${error.config?.url} (ServerTrace: ${serverTraceId})`,
      error.stack || ''
    );
    return Promise.reject(error);
  }
);

容灾回滚:前端静态资源 CDN 降级与 Feature Flag 快速切流。

拿到了完备的现场证据链,如果确认是前端新发布的 JS Bundle 资源在部分特定 iOS 版本下崩溃,接下来要做的是分钟级的止损与容灾。高并发业务切忌直接重新走漫长的 CI/CD 构建流水线。

可为静态资源准备故障域名和版本回退策略,并用 Feature Flag 控制可独立降级的功能。开关配置本身也要有缓存、签名、超时和默认行为;涉及支付等关键路径时,应先验证旧链路仍可用且状态可兼容。

# 在边缘 CDN 节点校验前端静态资源 Hash 一致性
curl -I -H "Accept-Encoding: gzip" https://cdn.internal.net/static/js/main.v2.4.1.js

# 查看边缘 Nginx 网关对前端 TraceHeader 的透传情况
kubectl logs -n prod-gateway -l app=ingress-envoy --since=15m | \
  grep "X-Client-Trace-Id" | awk '{print $1, $7, $9, $14}' | head -n 10

全局异常捕获、离线队列、请求关联 ID 和功能降级能缩短定位时间。离线日志要设置大小和保留上限,避免写入敏感输入;sendBeacon 也不是可靠投递,服务端仍要做采样和去重。

转载自 CSDN-专业IT技术社区

原文链接:https://blog.csdn.net/2609_95049439/article/details/163997427

文章来源转载

评论

赞0

评论列表

微信小程序
QQ小程序

关于作者

点赞数:0
关注数:0
粉丝:0
文章:0
关注标签:0
加入于:--