To find out if this integration is available in your organization, see your Datadog Integrations page or ask your organization administrator.

To initiate an exception request to enable this integration for your organization, email support@ddog-gov.com.

概要

Google Cloud TPU 製品は、スケーラブルで使いやすいクラウドコンピューティングリソースを通じて Tensor Processing Unit (TPU) を利用できるようにします。ML 研究者、ML エンジニア、開発者、データサイエンティストの誰もが最先端の ML (機械学習) モデルを実行できます。

Datadog Google Cloud Platform インテグレーションを使用して、Google Cloud TPU からメトリクスを収集できます。

セットアップ

インストール

Google Cloud TPU を使用するには、Google Cloud Platform インテグレーションを設定するだけです。

ログ収集

Google Cloud TPU のログは Google Cloud Logging で収集され、Cloud Pub/Sub トピックを通じて Dataflow ジョブに送信されます。まだの場合は、Datadog Dataflow テンプレートでロギングをセットアップしてください

これが完了したら、Google Cloud TPU のログを Google Cloud Logging から Pub/Sub トピックへエクスポートします。

  1. Google Cloud Logging のページに移動し、Google Cloud TPU のログを絞り込みます。
  2. Create Export をクリックし、シンクに名前を付けます。
  3. 宛先として “Cloud Pub/Sub” を選択し、その目的で作成された Pub/Sub トピックを選択します。: Pub/Sub トピックは別のプロジェクトに配置できます。
  4. 作成をクリックし、確認メッセージが表示されるまで待ちます。

収集データ

メトリクス

gcp.tpu.cpu.utilization
(gauge)
Current CPU utilization on the TPU worker, represented as a percentage. Values are typically numbers between 0.0 and 100.0, but might exceed 100.0.
Shown as percent
gcp.tpu.memory.usage
(gauge)
Memory usage in bytes.
Shown as byte
gcp.tpu.network.received_bytes_count
(count)
Cumulative bytes of data this server has received over the network.
Shown as byte
gcp.tpu.network.sent_bytes_count
(count)
Cumulative bytes of data this server has sent over the network.
Shown as byte
gcp.tpu.accelerator.duty_cycle
(count)
Percentage of time over the sample period during which the accelerator was actively processing
Shown as percent
gcp.tpu.instance.uptime_total
(count)
Elapsed time since the VM was started, in seconds.
Shown as second
gcp.gke.node.accelerator.tensorcore_utilization
(count)
Current percentage of the Tensorcore that is utilized.
Shown as percent
gcp.gke.node.accelerator.duty_cycle
(count)
Percent of time over the past sample period (10s) during which the accelerator was actively processing.
Shown as percent
gcp.gke.node.accelerator.memory_used
(count)
Total accelerator memory allocated in bytes.
Shown as byte
gcp.gke.node.accelerator.memory_total
(count)
Total accelerator memory in bytes.
Shown as byte
gcp.gke.node.accelerator.memory_bandwidth_utilization
(count)
Current percentage of the accelerator memory bandwidth that is being used.
Shown as percent
gcp.gke.container.accelerator.tensorcore_utilization
(count)
Current percentage of the Tensorcore that is utilized.
Shown as percent
gcp.gke.container.accelerator.duty_cycle
(count)
Percent of time over the past sample period (10s) during which the accelerator was actively processing.
Shown as percent
gcp.gke.container.accelerator.memory_used
(count)
Total accelerator memory allocated in bytes.
Shown as byte
gcp.gke.container.accelerator.memory_total
(count)
Total accelerator memory in bytes.
Shown as byte
gcp.gke.container.accelerator.memory_bandwidth_utilization
(count)
Current percentage of the accelerator memory bandwidth that is being used.
Shown as percent
gcp.tpu.accelerator.memory_bandwidth_utilization
(gauge)
Current percentage of the accelerator memory bandwidth that is being used. Computed by dividing the memory bandwidth used over a sample period by the maximum supported bandwidth over the same sample period.
Shown as percent
gcp.tpu.accelerator.memory_total
(gauge)
Total accelerator memory currently allocated in bytes.
Shown as byte
gcp.tpu.accelerator.memory_used
(gauge)
Total accelerator memory currently used in bytes.
Shown as byte
gcp.tpu.accelerator.tensorcore_utilization
(gauge)
Current percentage of the Tensorcore that is utilized. Computed by dividing the Tensorcore operations that were performed over a sample period by the supported number of Tensorcore operations over the same sample period.
Shown as percent
gcp.tpu.multislice.accelerator.device_to_host_transfer_latencies.avg
(gauge)
Cumulative distribution of device to host transfer latency for each chunk of data. A latency starts when the request for data to be transferred to the host is issued and ends when an acknowledgement is received that the transfer of data has completed.
Shown as microsecond
gcp.tpu.multislice.accelerator.host_to_device_transfer_latencies.avg
(gauge)
Cumulative distribution of host to device transfer latency for each chunk of data of multislice traffic. A latency starts when the request for data to be transferred to the device is issued and ends when an acknowledgement is received that the transfer of data has completed.
Shown as microsecond
gcp.tpu.multislice.network.collective_end_to_end_latencies.avg
(gauge)
Cumulative distribution of end to end collective latency for multislice traffic. A latency starts when the request for the collective is issued and ends when an acknowledgement is received that the transfer of data has completed.
Shown as microsecond
gcp.tpu.multislice.network.dcn_transfer_latencies.avg
(gauge)
Cumulative distribution of network-transfer latencies for multislice traffic. A latency starts when the request for data to be transferred over the DCN is issued and ends when an acknowledgement is received that the transfer of data has completed.
Shown as microsecond

イベント

Google Cloud TPU インテグレーションには、イベントは含まれません。

サービスチェック

Google Cloud TPU インテグレーションには、サービスのチェック機能は含まれません。

トラブルシューティング

ご不明な点は、Datadog のサポートチームまでお問い合わせください。