### Debugger V2 # Debugging Numerical Issues in TensorFlow Programs Using TensorBoard Debugger V2 > *NOTE*: tf.debugging.experimental.enable_dump_debug_info() is an experimental > API and may be subject to breaking changes in the future. Catastrophic events involving [NaN](https://en.wikipedia.org/wiki/NaN)s can sometimes occur during a TensorFlow program, crippling the model training processes. The root cause of such events are often obscure, especially for models of non-trivial size and complexity. To make it easier to debug this type of model bugs, TensorBoard 2.3+ (together with TensorFlow 2.3+) provides a specialized dashboard called Debugger V2. Here we demonstrate how to use this tool by working through a real bug involving NaNs in a neural network written in TensorFlow. The techniques illustrated in this tutorial are applicable to other types of debugging activities such as inspecting runtime tensor shapes in complex programs. This tutorial focuses on NaNs due to their relatively high frequency of occurrence. ## Observing the bug The source code of the TF2 program we’ll debug is [available on GitHub](https://github.com/tensorflow/tensorflow/blob/master/tensorflow/python/debug/examples/v2/debug_mnist_v2.py). The example program is also packaged into the tensorflow pip package (version 2.3+) and can be invoked by: ```sh python -m tensorflow.python.debug.examples.v2.debug_mnist_v2 ``` This TF2 program creates a multi-layer perception (MLP) and trains it to recognize [MNIST](https://en.wikipedia.org/wiki/MNIST_database) images. This example purposefully uses the low-level API of TF2 to define custom layer constructs, loss function, and training loop, because the likelihood of NaN bugs is higher when we use this more flexible but more error-prone API than when we use the easier-to-use but slightly less flexible high-level APIs such as [tf.keras](https://www.tensorflow.org/guide/keras). The program prints a test accuracy after each training step. We can see in the console that the test accuracy gets stuck at a near-chance level (~0.1) after the first step. This is certainly not how the model training is expected to behave: we expect the accuracy to gradually approach 1.0 (100%) as the step increases. ``` Accuracy at step 0: 0.216 Accuracy at step 1: 0.098 Accuracy at step 2: 0.098 Accuracy at step 3: 0.098 ... ``` An educated guess is that this problem is caused by a numerical instability, such as NaN or infinity. However, how do we confirm this is really the case and how do we find the TensorFlow operation (op) responsible for generating the numerical instability? To answer these questions, let’s instrument the buggy program with Debugger V2. ## Instrumenting TensorFlow code with Debugger V2 [`tf.debugging.experimental.enable_dump_debug_info()`](https://www.tensorflow.org/api_docs/python/tf/debugging/experimental/enable_dump_debug_info) is the API entry point of Debugger V2. It instruments a TF2 program with a single line of code. For instance, adding the following line near the beginning of the program will cause debug information to be written to the log directory (logdir) at /tmp/tfdbg2_logdir. The debug information covers various aspects of TensorFlow runtime. In TF2, it includes the full history of eager execution, graph building performed by [@tf.function](https://www.tensorflow.org/api_docs/python/tf/function), the execution of the graphs, the tensor values generated by the execution events, as well as the code location (Python stack traces) of those events. The richness of the debug information enables users to narrow in on obscure bugs. ```py tf.debugging.experimental.enable_dump_debug_info( "/tmp/tfdbg2_logdir", tensor_debug_mode="FULL_HEALTH", circular_buffer_size=-1) ``` The `tensor_debug_mode` argument controls what information Debugger V2 extracts from each eager or in-graph tensor. “FULL_HEALTH” is a mode that captures the following information about each floating-type tensor (e.g., the commonly-seen float32 and the less common [bfloat16](https://en.wikipedia.org/wiki/Bfloat16_floating-point_format) dtype): - DType - Rank - Total number of elements - A breakdown of the floating-type elements into the following categories: negative finite (`-`), zero (`0`), positive finite (`+`), negative infinity (`-∞`), positive infinity (`+∞`), and `NaN`. The “FULL_HEALTH” mode is suitable for debugging bugs involving NaN and infinity. See below for other supported `tensor_debug_mode`s. The `circular_buffer_size` argument controls how many tensor events are saved to the logdir. It defaults to 1000, which causes only the last 1000 tensors before the end of the instrumented TF2 program to be saved to disk. This default behavior reduces debugger overhead by sacrificing debug-data completeness. If the completeness is preferred, as in this case, we can disable the circular buffer by setting the argument to a negative value (e.g., -1 here). The debug_mnist_v2 example invokes `enable_dump_debug_info()` by passing command-line flags to it. To run our problematic TF2 program again with this debugging instrumentation enabled, do: ```sh python -m tensorflow.python.debug.examples.v2.debug_mnist_v2 \ --dump_dir /tmp/tfdbg2_logdir --dump_tensor_debug_mode FULL_HEALTH ``` ## Starting the Debugger V2 GUI in TensorBoard Running the program with the debugger instrumentation creates a logdir at /tmp/tfdbg2_logdir. We can start TensorBoard and point it at the logdir with: ```sh tensorboard --logdir /tmp/tfdbg2_logdir ``` In the web browser, navigate to TensorBoard’s page at http://localhost:6006. The “Debugger V2” plugin will be inactive by default, so select it from the “Inactive plugins” menu at top right. Once selected, it should look like the following: ## Using Debugger V2 GUI to find the root cause of NaNs The Debugger V2 GUI in TensorBoard is organized into six sections: - **Alerts**: This top-left section contains a list of “alert” events detected by the debugger in the debug data from the instrumented TensorFlow program. Each alert indicates a certain anomaly that warrants attention. In our case, this section highlights 499 NaN/∞ events with a salient pink-red color. This confirms our suspicion that the model fails to learn because of the presence of NaNs and/or infinities in its internal tensor values. We’ll delve into these alerts shortly. - **Python Execution Timeline**: This is the upper half of the top-middle section. It presents the full history of the eager execution of ops and graphs. Each box of the timeline is marked by the initial letter of the op or graph’s name (e.g., “T” for the “TensorSliceDataset” op, “m” for the “model” `tf.function`). We can navigate this timeline by using the navigation buttons and the scrollbar above the timeline. - **Graph Execution** : Located at the top-right corner of the GUI, this section will be central to our debugging task. It contains a history of all the floating-dtype tensors computed inside graphs (i.e., compiled by `@tf-function`s). - **Graph Structure** (bottom half of the top-middle section), **Source Code** (bottom-left section), and **Stack Trace** (bottom-right section) are initially empty. Their contents will be populated when we interact with the GUI. These three sections will also play important roles in our debugging task. Having oriented ourselves to the organization of the UI, let’s take the following steps to get to the bottom of why the NaNs appeared. First, click the **NaN/∞** alert in the Alerts section. This automatically scrolls the list of 600 graph tensors in the Graph Execution section and focuses on the #88, which is a tensor named `Log:0` generated by a `Log` (natural logarithm) op. A salient pink-red color highlights a -∞ element among the 1000 elements of the 2D float32 tensor. This is the first tensor in the TF2 program’s runtime history that contained any NaN or infinity: tensors computed before it do not contain NaN or ∞; many (in fact, most) tensors computed afterwards contain NaNs. We can confirm this by scrolling up and down the Graph Execution list. This observation provides a strong hint that the `Log` op is the source of the numerical instability in this TF2 program. Why does this `Log` op spit out a -∞? Answering that question requires examining the input to the op. Clicking on the name of the tensor (`Log:0`) brings up a simple but informative visualization of the `Log` op’s vicinity in its TensorFlow graph in the Graph Structure section. Note the top-to-bottom direction of information flow. The op itself is shown in the bold in the middle. Immediately above it, we can see a Placeholder op provides the one and only input to the `Log` op. Where is the tensor generated by this `probs` Placeholder in the Graph Execution list? By using the yellow background color as a visual aid, we can see that the `probs:0` tensor is three rows above the `Log:0` tensor, that is, in row 85. A more careful look at the numerical breakdown of the `probs:0` tensor in row 85 reveals why its consumer `Log:0` produces a -∞: Among the 1000 elements of `probs:0`, one element has a value of 0. The -∞ is a result of computing the natural logarithm of 0! If we can somehow ensure that the `Log` op gets exposed to only positive inputs, we’ll be able to prevent the NaN/∞ from happening. This can be achieved by applying clipping (e.g., by using [`tf.clip_by_value()`](https://www.tensorflow.org/api_docs/python/tf/clip_by_value)) on the Placeholder `probs` tensor. We are getting closer to solving the bug, but not quite done yet. In order to apply the fix, we need to know where in the Python source code the `Log` op and its Placeholder input originated. Debugger V2 provides first-class support for tracing the graph ops and execution events to their source. When we clicked the `Log:0` tensor in Graph Executions, the Stack Trace section was populated with the original stack trace of the `Log` op’s creation. The stack trace is somewhat large because it includes many frames from TensorFlow’s internal code (e.g., gen_math_ops.py and dumping_callback.py), which we can safely ignore for most debugging tasks. The frame of interest is Line 216 of debug_mnist_v2.py (i.e, the Python file we’re actually trying to debug). Clicking “Line 216” brings up a view of the corresponding line of code in the Source Code section. This finally brings us to the source code that created the problematic `Log` op from its `probs` input. This is our custom categorical cross-entropy loss function decorated with `@tf.function` and hence converted into a TensorFlow graph. The Placeholder op `probs` corresponds to the first input argument to the loss function. The `Log` op is created with the tf.math.log() API call. The value-clipping fix to this bug will look something like: ```py diff = -(labels * tf.math.log(tf.clip_by_value(probs), 1e-6, 1.)) ``` It will resolve the numerical instability in this TF2 program and cause the MLP to train successfully. Another possible approach to fixing the numerical instability is to use [`tf.keras.losses.CategoricalCrossentropy`](https://www.tensorflow.org/api_docs/python/tf/keras/losses/CategoricalCrossentropy). This concludes our journey from observing a TF2 model bug to coming up with a code change that fixes the bug, aided by the Debugger V2 tool, which provides full visibility into the eager and graph execution history of the instrumented TF2 program, including the numerical summaries of tensor values and association between ops, tensors and their original source code. ## Hardware compatibility of Debugger V2 Debugger V2 supports mainstream training hardware including CPU and GPU. Multi-GPU training with [tf.distributed.MirroredStrategy](https://www.tensorflow.org/api_docs/python/tf/distribute/MirroredStrategy) is also supported. The support for [TPU](https://www.tensorflow.org/guide/tpu) is still in an early stage and requires calling ```py tf.config.set_soft_device_placement(True) ``` before calling `enable_dump_debug_info()`. It may have other limitations on TPUs as well. If you run into problems using Debugger V2, please report bugs on our [GitHub issues page](https://github.com/tensorflow/tensorboard/issues). ## API compatibility of Debugger V2 Debugger V2 is implemented at a relatively low level of TensorFlow’s software stack, and hence is compatible with [tf.keras](https://www.tensorflow.org/api_docs/python/tf/keras), [tf.data](https://www.tensorflow.org/guide/data), and other APIs built on top of TensorFlow’s lower levels. Debugger V2 is also backward compatible with TF1, although the Eager Execution Timeline will be empty for the debug logdirs generated by TF1 programs. ## API usage tips A frequently-asked question about this debugging API is where in the TensorFlow code one should insert the call to `enable_dump_debug_info()`. Typically, the API should be called as early as possible in your TF2 program, preferably after the Python import lines and before graph building and execution begin. This will ensure full coverage of all the ops and graphs that power your model and its training. The currently supported tensor_debug_modes are: `NO_TENSOR`, `CURT_HEALTH`, `CONCISE_HEALTH`, `FULL_HEALTH`, and `SHAPE`. They vary in the amount of information extracted from each tensor and the performance overhead to the debugged program. Please refer to the [args section](https://www.tensorflow.org/api_docs/python/tf/debugging/experimental/enable_dump_debug_info) of `enable_dump_debug_info()`’s documentation. ## Performance overhead The debugging API introduces performance overhead to the instrumented TensorFlow program. The overhead varies by `tensor_debug_mode`, hardware type, and nature of the instrumented TensorFlow program. As a reference point, on a GPU, the `NO_TENSOR` mode adds a 15% overhead during the training of a [Transformer model](https://github.com/tensorflow/models/tree/master/official/legacy/transformer) under batch size 64. The percent overhead for other tensor_debug_modes are higher: approximately 50% for the `CURT_HEALTH`, `CONCISE_HEALTH`, `FULL_HEALTH` and `SHAPE` modes. On CPUs, the overhead is slightly lower. On TPUs, the overhead is currently higher. ## Relation to other TensorFlow debugging APIs Note that TensorFlow offers other tools and APIs for debugging. You can browse such APIs under the [`tf.debugging.*` namespace](https://www.tensorflow.org/api_docs/python/tf/debugging) at the API docs page. Among these APIs the most frequently used is [`tf.print()`](https://www.tensorflow.org/api_docs/python/tf/print). When should one use Debugger V2 and when should `tf.print()` be used instead? `tf.print()` is convenient in case where 1. we know exactly which tensors to print, 2. we know where exactly in the source code to insert those `tf.print()` statements, 3. the number of such tensors is not too large. For other cases (e.g., examining many tensor values, examining tensor values generated by TensorFlow’s internal code, and searching for the origin of numerical instability as we showed above), Debugger V2 provides a faster way of debugging. In addition, Debugger V2 provides a unified approach to inspecting eager and graph tensors. It additionally provides information about graph structure and code locations, which are beyond the capability of `tf.print()`. Another API that can be used to debug issues involving ∞ and NaN is [`tf.debugging.enable_check_numerics()`](https://www.tensorflow.org/api_docs/python/tf/debugging/enable_check_numerics). Unlike `enable_dump_debug_info()`, `enable_check_numerics()` does not save debug information on the disk. Instead, it merely monitors ∞ and NaN during TensorFlow runtime and errors out with the origin code location as soon as any op generates such bad numerical values. It has a lower performance overhead compared to `enable_dump_debug_info()`, but doesn’t afford a full trace of the program’s execution history and does not come with a graphical user interface like Debugger V2. --- ### What If Tool # Model Understanding with the What-If Tool Dashboard > **Warning** > This documentation only applies to TensorBoard 2.11 and earlier, as the > What-If Tool is no longer actively maintained. Please check out the actively > maintained [Learning Interpretability Tool > (LIT)](https://pair-code.github.io/lit/) instead. The What-If Tool (WIT) provides an easy-to-use interface for expanding understanding of black-box classification and regression ML models. With the plugin, you can perform inference on a large set of examples and immediately visualize the results in a variety of ways. Additionally, examples can be edited manually or programmatically and re-run through the model in order to see the results of the changes. It contains tooling for investigating model performance and fairness over subsets of a dataset. The purpose of the tool is to give people a simple, intuitive, and powerful way to explore and investigate trained ML models through a visual interface with absolutely no code required. The tool can be accessed through TensorBoard or directly in a Jupyter or Colab notebook. For more in-depth details, demos, walkthroughs, and information specific to using WIT in notebook mode, see the [What-If Tool website](https://pair-code.github.io/what-if-tool). ## Requirements To use WIT in TensorBoard, two things are necessary: - The model(s) you wish to explore must be served using [TensorFlow Serving](https://github.com/tensorflow/serving) using the classify, regress, or predict API. - The dataset to be inferred by the models must be in a TFRecord file accessible by the TensorBoard web server. ## Usage When opening the What-If Tool dashboard in TensorBoard, you will see a setup screen where you provide the host and port of the model server, the name of the model being served, the type of model, and the path to the TFRecords file to load. After filling this information out and clicking "Accept", WIT will load the dataset and run inference with the model, displaying the results. For details on the different features of WIT and how they can aid in model understanding and fairness investigations, see the walkthrough on the [What-If Tool website](https://pair-code.github.io/what-if-tool). ## Demo model and dataset If you want to test out WIT in TensorBoard with a pre-trained model, you can download and unzip a pre-trained model and dataset from https://storage.googleapis.com/what-if-tool-resources/uci-census-demo/uci-census-demo.zip. The model is a binary classification model that uses the [UCI Census](https://archive.ics.uci.edu/ml/datasets/census+income) dataset to predict whether a person earns more than $50k a year. This dataset and prediction task is often used in machine learning modeling and fairness research. Set the environment variable MODEL_PATH to the location of the resulting model directory on your machine. Install docker and TensorFlow Serving following the [official documentation](https://www.tensorflow.org/tfx/serving/docker). Serve the model using docker through `docker run -p 8500:8500 --mount type=bind,source=${MODEL_PATH},target=/models/uci_income -e MODEL_NAME=uci_income -t tensorflow/serving`. Note you may need to run the command with `sudo` depending on your docker setup. Now launch tensorboard and use the dashboard drop-down to navigate to the What-If Tool. On the setup screen, set the inference adddress to "localhost:8500", the model name to "uci_income" and the path to examples to the full path to the downloaded `adult.tfrecord` file, then press "Accept". Some things to try with the What-If Tool on this demo include: - Editing a single datapoint and seeing the resulting change in inference. - Exploring the relationship between individual features in the dataset and the model's inference results through partial dependence plots. - Slicing the dataset into subsets and comparing the performance between slices. For an in-depth look at tool's features, check out the [What-If Tool walkthrough](https://pair-code.github.io/what-if-tool/walkthrough.html). Note the ground truth feature in the dataset that this model is trying to predict is named "Target", so when using the "Performance & Fairness" tab, "Target" is what you will want to specify in the ground truth feature dropdown. --- ### CONTRIBUTING ### Contributor License Agreements We'd love to accept your patches! Before we can take them, we have to jump a couple of legal hurdles. Please fill out either the individual or corporate Contributor License Agreement (CLA). * If you are an individual writing original source code and you're sure you own the intellectual property, then you'll need to sign an [individual CLA](http://code.google.com/legal/individual-cla-v1.0.html). * If you work for a company that wants to allow you to contribute your work, then you'll need to sign a [corporate CLA](http://code.google.com/legal/corporate-cla-v1.0.html). Follow either of the two links above to access the appropriate CLA and instructions for how to sign and return it. Once we receive it, we'll be able to accept your pull requests. ***NOTE***: Only original source code from you and other people that have signed the CLA can be accepted into the main repository. ### Working with the team If you're planning a larger contribution, please get in touch with the team through a GitHub issue before starting work - we can help guide you, and coordinating up front will make the process smoother. If you want to add a major feature, it may be a good candidate for adding a plugin. Let us know via a GitHub issue, and we can guide you in the process. ### Code reviews All submissions, including submissions by project members, require review. We use GitHub pull requests for this purpose. --- ### SECURITY # TensorBoard Security Please refer to [TensorFlow’s security model and guidelines][tf-security]. [tf-security]: https://github.com/tensorflow/tensorflow/blob/master/SECURITY.md ---