README (README.md)
<div align="center" markdown="1">
<img src="documentation/assets/SG_img/SG - Horizontal Glow 2.png" width="600"/>
<br/><br/>
Build, train, and fine-tune production-ready deep learning SOTA vision models
[](https://twitter.com/intent/tweet?text=Easily%20train%20or%20fine-tune%20SOTA%20computer%20vision%20models%20from%20one%20training%20repository&url=https://github.com/Deci-AI/super-gradients&via=deci_ai&hashtags=AI,deeplearning,computervision,training,opensource)
#### Version 3.5 is out! Notebooks have been updated!
______________________________________________________________________
</div>
<div align="center">
<p align="center">
<a href="#getting-started">Getting Started</a> β’
<a href="#implemented-model-architectures">Pretrained Models</a> β’
<a href="#community">Community</a> β’
<a href="#license">License</a> β’
</p>
<p align="center">
<a href="https://github.com/Deci-AI/super-gradients#prerequisites"><img src="https://img.shields.io/badge/python-3.7%20%7C%203.8%20%7C%203.9-blue" /></a>
<a href="https://github.com/Deci-AI/super-gradients#prerequisites"><img src="https://img.shields.io/badge/pytorch-1.9%20%7C%201.10-blue" /></a>
<a href="https://pypi.org/project/super-gradients/"><img src="https://img.shields.io/pypi/v/super-gradients" /></a>
<a href="https://github.com/Deci-AI/super-gradients/blob/master/documentation/source/model_zoo.md" ><img src="https://img.shields.io/badge/pre--trained%20models-34-brightgreen" /></a>
<a href="https://github.com/Deci-AI/super-gradients/releases"><img src="https://img.shields.io/github/v/release/Deci-AI/super-gradients" /></a>
<a href="https://join.slack.com/t/supergradients-comm52/shared_invite/zt-10vz6o1ia-b_0W5jEPEnuHXm087K~t8Q"><img src="https://img.shields.io/badge/slack-community-blueviolet" /></a>
<a href="https://github.com/Deci-AI/super-gradients/blob/master/LICENSE.md"><img src="https://img.shields.io/badge/license-Apache%202.0-blue" /></a>
<a href="https://docs.deci.ai/super-gradients/documentation/source/welcome.html"><img src="https://img.shields.io/badge/docs-mkdocs-brightgreen" /></a>
</p>
</div>
______________________________________________________________________
Build with SuperGradients
__________________________________________________________________________________________________________
Support various computer vision tasks
<div align="center">
<img src="https://github.com/Deci-AI/super-gradients/raw/master/documentation/assets/SG_img/Segmentation 1500x900 .png" width="250px">
<img src="https://github.com/Deci-AI/super-gradients/raw/master/documentation/assets/SG_img/Object detection 1500X900.png" width="250px">
<img src="https://github.com/Deci-AI/super-gradients/raw/master/documentation/assets/SG_img/Classification 1500x900.png" width="250px">
<img src="https://github.com/Deci-AI/super-gradients/raw/master/documentation/assets/SG_img/PoseEstimation.jpg" width="250px">
</div>
Ready to deploy pre-trained SOTA models
YOLO-NAS and YOLO-NAS-POSE architectures are out!
The new YOLO-NAS delivers state-of-the-art performance with the unparalleled accuracy-speed performance, outperforming other models such as YOLOv5, YOLOv6, YOLOv7 and YOLOv8.
A YOLO-NAS-POSE model for pose estimation is also available, delivering state-of-the-art accuracy/performance tradeoff.
Check these out here: YOLO-NAS & YOLO-NAS-POSE.
<div align="center">
<img src="https://github.com/Deci-AI/super-gradients/raw/master/documentation/source/images/yolo_nas_frontier.png" height="600px">
<img src="https://github.com/Deci-AI/super-gradients/raw/master/documentation/source/images/yolo_nas_pose_frontier_t4.png" height="600px">
</div>
Load model with pretrained weights
from super_gradients.training import models
from super_gradients.common.object_names import Modelsmodel = models.get(Models.YOLO_NAS_M, pretrained_weights="coco")
#### All Computer Vision Models - Pretrained Checkpoints can be found in the Model Zoo
#### Classification
<div align="center">
<img src="https://github.com/Deci-AI/super-gradients/raw/master/documentation/assets/SG_img/[email protected]" width="800px">
</div>
#### Semantic Segmentation
<div align="center">
<img src="https://github.com/Deci-AI/super-gradients/raw/master/documentation/assets/SG_img/Semantic [email protected]" width="800px">
</div>
#### Object Detection
<div align="center">
<img src="https://github.com/Deci-AI/super-gradients/raw/master/documentation/assets/SG_img/Object [email protected]" width="800px">
</div>
#### Pose Estimation
<div align="center">
<img src="https://github.com/Deci-AI/super-gradients/raw/master/documentation/source/images/yolo_nas_pose_frontier_t4.png" width="400px">
<img src="https://github.com/Deci-AI/super-gradients/raw/master/documentation/source/images/yolo_nas_pose_frontier_xavier_nx.png" width="400px">
</div>
Easy to train SOTA Models
Easily load and fine-tune production-ready, pre-trained SOTA models that incorporate best practices and validated hyper-parameters for achieving best-in-class accuracy.
For more information on how to do it go to Getting Started
#### Plug and play recipes
python -m super_gradients.train_from_recipe architecture=regnetY800 dataset_interface.data_dir=<YOUR_Imagenet_LOCAL_PATH> ckpt_root_dir=<CHEKPOINT_DIRECTORY>More examples on how and why to use recipes can be found in Recipes
Production readiness
All SuperGradients modelsβ are production ready in the sense that they are compatible with deployment tools such as TensorRT (Nvidia) and OpenVINO (Intel) and can be easily taken into production. With a few lines of code you can easily integrate the models into your codebase.
Load model with pretrained weights
from super_gradients.training import models
from super_gradients.common.object_names import Modelsmodel = models.get(Models.YOLO_NAS_M, pretrained_weights="coco")
Prepare model for conversion
Input size is in format of [Batch x Channels x Width x Height] where 640 is the standard COCO dataset dimensions
model.eval()
model.prep_model_for_conversion(input_size=[1, 3, 640, 640])
Create dummy_input
Convert model to onnx
torch.onnx.export(model, dummy_input, "yolo_nas_m.onnx")More information on how to take your model to production can be found in Getting Started notebooks
Quick Installation
__________________________________________________________________________________________________________
pip install super-gradientsWhat's New
__________________________________________________________________________________________________________
Version 3.4.0 (November 6, 2023)
* YoloNAS-Pose model released - a new frontier in pose estimation
* Added option to export a recipe to a single YAML file or to a standalone train.py file
* Other bugfixes & minor improvements. Full release notes available here
__________________________________________________________________________________________________________
Version 3.1.3 (July 19, 2023)
* Pose Estimation Task Support - Check out fine-tuning notebook example
* Pre-trained modified DEKR model for pose estimation (TensorRT-compatible)
* Support for Python 3.10
* Support for torch.compile
* Other bugfixes & minor improvements. Check out release notes
__________________________________________________________________________________________________________
30th of May
* Quantization Aware Training YoloNAS on Custom Dataset
Version 3.1.1 (May 3rd)
* YOLO-NAS
* New predict function (predict on any image, video, url, path, stream)
* RoboFlow100 datasets integration
* A new Documentation Hub
* Integration with DagsHub for experiment monitoring
* Support Darknet/Yolo format detection dataset (used by Yolo v5, v6, v7, v8)
* Segformer model and recipe
* Post Training Quantization and Quantization Aware Training - notebooks
Check out SG full release notes.
Table of Content
__________________________________________________________________________________________________________
- Getting Started
- Advanced Features
- Installation Methods
- Prerequisites
- Quick Installation
- Implemented Model Architectures
- Contributing
- Citation
- Community
- License
- Deci Platform
Getting Started
__________________________________________________________________________________________________________
Start Training with Just 1 Command Line
The most simple and straightforward way to start training SOTA performance models with SuperGradients reproducible recipes. Just define your dataset path and where you want your checkpoints to be saved and you are good to go from your terminal!
Just make sure that you setup your dataset according to the data dir specified in the recipe.
python -m super_gradients.train_from_recipe --config-name=imagenet_regnetY architecture=regnetY800 dataset_interface.data_dir=<YOUR_Imagenet_LOCAL_PATH> ckpt_root_dir=<CHEKPOINT_DIRECTORY>Quickly Load Pre-Trained Weights for Your Desired Model with SOTA Performance
Want to try our pre-trained models on your machine? Import SuperGradients, initialize your Trainer, and load your desired architecture and pre-trained weights from our SOTA model zoo
The pretrained_weights argument will load a pre-trained architecture on the provided dataset
import super_gradientsmodel = models.get("model-name", pretrained_weights="pretrained-model-name")
Classification
* [](https://colab.research.google.com/github/Deci-AI/super-gradients/blob/master/notebooks/transfer_learning_classification.ipynb) Transfer Learning for classification
* [](https://colab.research.google.com/github/Deci-AI/super-gradients/blob/master/notebooks/PTQ_and_QAT_for_classification.ipynb) PTQ and QAT for classification
Semantic Segmentation
* [](https://colab.research.google.com/github/Deci-AI/super-gradients/blob/master/notebooks/quickstart_segmentation.ipynb) Segmentation Quick Start
* [](https://colab.research.google.com/github/Deci-AI/super-gradients/blob/master/notebooks/transfer_learning_semantic_segmentation.ipynb) Segmentation Transfer Learning
* [](https://colab.research.google.com/github/Deci-AI/super-gradients/blob/master/notebooks/segmentation_connect_custom_dataset.ipynb) How to Connect Custom Dataset
* [](https://colab.research.google.com/github/Deci-AI/super-gradients/blob/master/notebooks/Segmentation_Model_Export.ipynb) How to export segmentation model to ONNX
Pose Estimation
* [](https://colab.research.google.com/github/Deci-AI/super-gradients/blob/master/notebooks/YoloNAS_Pose_Fine_Tuning_Animals_Pose_Dataset.ipynb) Fine Tuning YoloNAS-Pose on AnimalPose dataset
* [](https://colab.research.google.com/github/Deci-AI/super-gradients/blob/master/notebooks/DEKR_PoseEstimationFineTuning.ipynb) Fine Tuning DEKR on AnimalPose dataset
Object Detection
* [](https://colab.research.google.com/github/Deci-AI/super-gradients/blob/master/notebooks/YoloNAS_Inference_using_TensorRT.ipynb) YoloNAS inference using TensorRT
* [](https://colab.research.google.com/github/Deci-AI/super-gradients/blob/master/notebooks/detection_transfer_learning.ipynb) Object Detection Transfer Learning
* [](https://colab.research.google.com/github/Deci-AI/super-gradients/blob/master/notebooks/detection_how_to_connect_custom_dataset.ipynb) How to Connect Custom Dataset
* [](https://colab.research.google.com/github/Deci-AI/super-gradients/blob/master/notebooks/yolo_nas_custom_dataset_fine_tuning_with_qat.ipynb) Quantization Aware Training YoloNAS on Custom Dataset
How to Predict Using Pre-trained Model
* [](https://colab.research.google.com/github/Deci-AI/super-gradients/blob/master/notebooks/how_to_run_model_predict.ipynb) How to Predict Using Pre-trained Model
Albumentations Integration
* [](https://colab.research.google.com/github/Deci-AI/super-gradients/blob/master/notebooks/albumentations_tutorial.ipynb) Using Albumentations with SG
Advanced Features
__________________________________________________________________________________________________________
Post Training Quantization and Quantization Aware Training
Quantization involves representing weights and biases in lower precision, resulting in reduced memory and computational requirements, making it useful for deploying models on devices with limited resources.
The process can be done during training, called Quantization aware training, or after training, called post-training quantization.
A full tutorial can be found here.
* [](https://bit.ly/3KrN6an) Post Training Quantization and Quantization Aware Training
Quantization Aware Training YoloNAS on Custom Dataset
This tutorial provides a comprehensive guide on how to fine-tune a YoloNAS model using a custom dataset.
It also demonstrates how to utilize SG's QAT (Quantization-Aware Training) support. Additionally, it offers step-by-step instructions on deploying the model and performing benchmarking.
* [](https://bit.ly/3MIKdTy) Quantization Aware Training YoloNAS on Custom Dataset
Knowledge Distillation Training
Knowledge Distillation is a training technique that uses a large model, teacher model, to improve the performance of a smaller model, the student model.
Learn more about SuperGradients knowledge distillation training with our pre-trained BEiT base teacher model and Resnet18 student model on CIFAR10 example notebook on Google Colab for an easy to use tutorial using free GPU hardware
* [](https://colab.research.google.com/github/Deci-AI/super-gradients/blob/master/notebooks/how_to_use_knowledge_distillation_for_classification.ipynb) Knowledge Distillation Training
Recipes
To train a model, it is necessary to configure 4 main components.
These components are aggregated into a single "main" recipe .yaml file that inherits the aforementioned dataset, architecture, raining and checkpoint params.
It is also possible (and recommended for flexibility) to override default settings with custom ones.
All recipes can be found here
</br>
Recipes support out of the box every model, metric or loss that is implemented in SuperGradients, but you can easily extend this to any custom object that you need by "registering it". Check out this tutorial for more information.
* [](https://colab.research.google.com/github/Deci-AI/super-gradients/blob/master/notebooks/what_are_recipes_and_how_to_use.ipynb) How to Use Recipes
<details markdown="1">
<summary><h3>Using Distributed Data Parallel (DDP) </h3></summary>
#### Why use DDP ?
Recent Deep Learning models are growing larger and larger to an extent that training on a single GPU can take weeks.
In order to train models in a timely fashion, it is necessary to train them with multiple GPUs.
Using 100s GPUs can reduce training time of a model from a week to less than an hour.
#### How does it work ?
Each GPU has its own process, which controls a copy of the model and which loads its own mini-batch from disk and sends
it to its GPU during training. After the forward pass is completed on every GPU, the gradient is reduced across all
GPUs, yielding to all the GPUs having the same gradient locally. This leads to the model weights to stay synchronized
across all GPUs after the backward pass.
#### How to use it ?
You can use SuperGradients to train your model with DDP in just a few lines.
main.py
from super_gradients import init_trainer, Trainer
from super_gradients.common import MultiGPUMode
from super_gradients.training.utils.distributed_training_utils import setup_deviceInitialize the environment
init_trainer()Launch DDP on 4 GPUs'
setup_device(multi_gpu=MultiGPUMode.DISTRIBUTED_DATA_PARALLEL, num_gpus=4)Call the trainer
Trainer(expriment_name=...)Everything you do below will run on 4 gpus
...
Trainer.train(...)
Finally, you can launch your distributed training with a simple python call.
python main.py
Please note that if you work with torch<1.9.0 (deprecated), you will have to launch your training with either torch.distributed.launch or torchrun, in which case nproc_per_node will overwrite the value set with gpu_mode:
python -m torch.distributed.launch --nproc_per_node=4 main.pytorchrun --nproc_per_node=4 main.py#### Calling functions on a single node
It is often in DDP training that we want to execute code on the master rank (i.e rank 0).
In SG, users usually execute their own code by triggering "Phase Callbacks" (see "Using phase callbacks" section below).
One can make sure the desired code will only be ran on rank 0, using ddp_silent_mode or the multi_process_safe decorator.
For example, consider the simple phase callback below, that uploads the first 3 images of every batch during training to
the Tensorboard:
from super_gradients.training.utils.callbacks import PhaseCallback, PhaseContext, Phase
from super_gradients.common.environment.env_helpers import multi_process_safeclass Upload3TrainImagesCalbback(PhaseCallback):
def __init__(
self,
):
super().__init__(phase=Phase.TRAIN_BATCH_END)
@multi_process_safe
def __call__(self, context: PhaseContext):
batch_imgs = context.inputs.cpu().detach().numpy()
tag = "batch_" + str(context.batch_idx) + "_images"
context.sg_logger.add_images(tag=tag, images=batch_imgs[: 3], global_step=context.epoch)
The @multi_process_safe decorator ensures that the callback will only be triggered by rank 0. Alternatively, this can also
be done by the SG trainer boolean attribute (which the phase context has access to), ddp_silent_mode, which is set to False
iff the current process rank is zero (even after the process group has been killed):
from super_gradients.training.utils.callbacks import PhaseCallback, PhaseContext, Phaseclass Upload3TrainImagesCalbback(PhaseCallback):
def __init__(
self,
):
super().__init__(phase=Phase.TRAIN_BATCH_END)
def __call__(self, context: PhaseContext):
if not context.ddp_silent_mode:
batch_imgs = context.inputs.cpu().detach().numpy()
tag = "batch_" + str(context.batch_idx) + "_images"
context.sg_logger.add_images(tag=tag, images=batch_imgs[: 3], global_step=context.epoch)
Note that ddp_silent_mode can be accessed through SgTrainer.ddp_silent_mode. Hence, it can be used in scripts after calling
SgTrainer.train() when some part of it should be ran on rank 0 only.
#### Good to know
Your total batch size will be (number of gpus x batch size), so you might want to increase your learning rate.
There is no clear rule, but a rule of thumb seems to be to linearly increase the learning rate with the number of gpus
</details>
<details markdown="1">
<summary><h3> Easily change architectures parameters </h3></summary>
from super_gradients.training import modelsinstantiate default pretrained resnet18
default_resnet18 = models.get(model_name="resnet18", num_classes=100, pretrained_weights="imagenet")instantiate pretrained resnet18, turning DropPath on with probability 0.5
droppath_resnet18 = models.get(model_name="resnet18", arch_params={"droppath_prob": 0.5}, num_classes=100, pretrained_weights="imagenet")instantiate pretrained resnet18, without classifier head. Output will be from the last stage before global pooling
backbone_resnet18 = models.get(model_name="resnet18", arch_params={"backbone_mode": True}, pretrained_weights="imagenet")</details>
<details markdown="1">
<summary><h3> Using phase callbacks </h3></summary>
from super_gradients import Trainer
from torch.optim.lr_scheduler import ReduceLROnPlateau
from super_gradients.training.utils.callbacks import Phase, LRSchedulerCallback
from super_gradients.training.metrics.classification_metrics import Accuracydefine PyTorch train and validation loaders and optimizer
define what to be called in the callback
rop_lr_scheduler = ReduceLROnPlateau(optimizer, mode="max", patience=10, verbose=True)define phase callbacks, they will fire as defined in Phase
phase_callbacks = [LRSchedulerCallback(scheduler=rop_lr_scheduler,
phase=Phase.VALIDATION_EPOCH_END,
metric_name="Accuracy")]create a trainer object, look the declaration for more parameters
trainer = Trainer("experiment_name")define phase_callbacks as part of the training parameters
train_params = {"phase_callbacks": phase_callbacks}</details>
<details markdown="1">
<summary><h3> Integration to DagsHub </h3></summary>
[](https://colab.research.google.com/drive/11fW56pMpwOMHQSbQW6xxMRYvw1mEC-t-?usp=sharing)
from super_gradients import Trainertrainer = Trainer("experiment_name")
model = ...
training_params = { ... # Your training params
"sg_logger": "dagshub_sg_logger", # DagsHub Logger, see class super_gradients.common.sg_loggers.dagshub_sg_logger.DagsHubSGLogger for details
"sg_logger_params": # Params that will be passes to __init__ of the logger super_gradients.common.sg_loggers.dagshub_sg_logger.DagsHubSGLogger
{
"dagshub_repository": "<REPO_OWNER>/<REPO_NAME>", # Optional: Your DagsHub project name, consisting of the owner name, followed by '/', and the repo name. If this is left empty, you'll be prompted in your run to fill it in manually.
"log_mlflow_only": False, # Optional: Change to true to bypass logging to DVC, and log all artifacts only to MLflow
"save_checkpoints_remote": True,
"save_tensorboard_remote": True,
"save_logs_remote": True,
}
}
</details>
<details>
<summary><h3> Integration to Weights and Biases </h3></summary>
from super_gradients import Trainercreate a trainer object, look the declaration for more parameters
trainer = Trainer("experiment_name")train_params = { ... # training parameters
"sg_logger": "wandb_sg_logger", # Weights&Biases Logger, see class WandBSGLogger for details
"sg_logger_params": # paramenters that will be passes to __init__ of the logger
{
"project_name": "project_name", # W&B project name
"save_checkpoints_remote": True
"save_tensorboard_remote": True
"save_logs_remote": True
}
}
</details>
<details markdown="1">
<summary><h3> Integration to ClearML </h3></summary>
from super_gradients import Trainercreate a trainer object, look the declaration for more parameters
trainer = Trainer("experiment_name")train_params = { ... # training parameters
"sg_logger": "clearml_sg_logger", # ClearML Logger, see class ClearMLSGLogger for details
"sg_logger_params": # paramenters that will be passes to __init__ of the logger
{
"project_name": "project_name", # ClearML project name
"save_checkpoints_remote": True,
"save_tensorboard_remote": True,
"save_logs_remote": True,
}
}
</details>
<details markdown="1">
<summary><h3> Integration to Voxel51 </h3></summary>
You can apply SuperGradients YOLO-NAS models directly to your FiftyOne dataset using the apply_model() method:
import fiftyone as fo
import fiftyone.zoo as fozfrom super_gradients.training import models
dataset = foz.load_zoo_dataset("quickstart", max_samples=25)
dataset.select_fields().keep_fields()
model = models.get("yolo_nas_m", pretrained_weights="coco")
dataset.apply_model(model, label_field="yolo_nas", confidence_thresh=0.7)
session = fo.launch_app(dataset)
The SuperGradients YOLO-NAS model can be accessed directly from the FiftyOne Model Zoo:
import fiftyone as fo
import fiftyone.zoo as fozmodel = foz.load_zoo_model("yolo-nas-torch")
dataset = foz.load_zoo_dataset("quickstart")
dataset.apply_model(model, label_field="yolo_nas")
session = fo.launch_app(dataset)
</details>
Installation Methods
__________________________________________________________________________________________________________
Prerequisites
<details markdown="1">
<summary>General requirements</summary>
- Python 3.7, 3.8 or 3.9 installed.
- 1.9.0 <= torch < 1.14
- https://pytorch.org/get-started/locally/
- The python packages that are specified in requirements.txt;
</details>
<details markdown="1">
<summary>To train on nvidia GPUs</summary>
- Nvidia CUDA Toolkit >= 11.2
- CuDNN >= 8.1.x
- Nvidia Driver with CUDA >= 11.2 support (β₯460.x)
</details>
Quick Installation
<details markdown="1">
<summary>Install stable version using PyPi</summary>
See in PyPi
pip install super-gradientsThat's it !
</details>
<details markdown="1">
<summary>Install using GitHub</summary>
pip install git+https://github.com/Deci-AI/super-gradients.git@stable</details>
Implemented Model Architectures
__________________________________________________________________________________________________________
All Computer Vision Models - Pretrained Checkpoints can be found in the Model Zoo
Image Classification
- DensNet (Densely Connected Convolutional Networks)
- DPN
- EfficientNet
- LeNet
- MobileNet
- MobileNet v2
- MobileNet v3
- PNASNet
- Pre-activation ResNet
- RegNet
- RepVGG
- ResNet
- ResNeXt
- SENet
- ShuffleNet
- ShuffleNet v2
- VGG
Semantic Segmentation
- PP-LiteSeg
- DDRNet (Deep Dual-resolution Networks)
- LadderNet
- RegSeg
- ShelfNet
- STDC
Object Detection
- CSP DarkNet
- DarkNet-53
- SSD (Single Shot Detector)
- YOLOX
Pose Estimation
__________________________________________________________________________________________________________
Implemented Datasets
__________________________________________________________________________________________________________
Deci provides implementation for various datasets. If you need to download any of the dataset, you can
find instructions.
Image Classification
- Cifar10
- ImageNet
Semantic Segmentation
- Cityscapes
- Coco
- PascalVOC 2012 / PascalAUG 2012
- SuperviselyPersons
- Mapillary Vistas Dataset
Object Detection
- Coco
- PascalVOC 2007 & 2012
Pose Estimation
- COCO
__________________________________________________________________________________________________________
Documentation
Check SuperGradients Docs for full documentation, user guide, and examples.
Contributing
To learn about making a contribution to SuperGradients, please see our Contribution page.
Our awesome contributors:
<a href="https://github.com/Deci-AI/super-gradients/graphs/contributors">
<img src="https://contrib.rocks/image?repo=Deci-AI/super-gradients" />
</a>
<br/>Made with contrib.rocks.
Citation
If you are using SuperGradients library or benchmarks in your research, please cite SuperGradients deep learning training library.
Community
If you want to be a part of SuperGradients growing community, hear about all the exciting news and updates, need help, request for advanced features,
or want to file a bug or issue report, we would love to welcome you aboard!
* Discord is the place to be and ask questions about SuperGradients and get support. Click here to join our Discord Community
* To report a bug, file an issue on GitHub.
* Join the SG Newsletter
for staying up to date with new features and models, important announcements, and upcoming events.
* For a short meeting with us, use this link and choose your preferred time.
License
This project is released under the Apache 2.0 license.
Citing
BibTeX
@misc{supergradients,
doi = {10.5281/ZENODO.7789328},
url = {https://zenodo.org/record/7789328},
author = {Aharon, Shay and {Louis-Dupont} and {Ofri Masad} and Yurkova, Kate and {Lotem Fridman} and {Lkdci} and Khvedchenya, Eugene and Rubin, Ran and Bagrov, Natan and Tymchenko, Borys and Keren, Tomer and Zhilko, Alexander and {Eran-Deci}},
title = {Super-Gradients},
publisher = {GitHub},
journal = {GitHub repository},
year = {2021},
}Latest DOI
[](https://doi.org/10.5281/zenodo.7789328)
__________________________________________________________________________________________________________
ΦΏ
Request free trial here
---
CONTRIBUTING (CONTRIBUTING.md)
Contribution Guidelines
Here is a simple guideline to get you started with your first contribution.
1. Set up your environment to follow our formatting guidelines and to use signed-commits.
2. Use issues to discuss the suggested changes. Create an issue describing changes if necessary and add labels to ease orientation.
3. Fork super-gradients so you can make local changes and test them.
4. Create a new branch for the issue. The branch naming convention is enforced by the CI/CD so please make sure you are using feature/SG- or hotfix/SG- format otherwise it will fail.
5. Implement your changes along with relevant tests for the issue. Please make sure you are covering unit, integration and e2e tests where required.
6. Create a pull request against master branch.
Code Style
We follow the reStructuredText docstring format (default of PyCharm), along with typing.
def python_function(first_argument: int, second_argument: int) -> str:
"""Do something with the two arguments. :param first_argument: First argument to the function
:param second_argument: Second argument to the function
:return: Description of the output
"""
Code Formatting
We enforce black code formatting in addition to existing flake8 checks.
To ensure everyone uses same code style, a project-wise configuration file has been added to SG repo. It ensures all formatting will be exactly the same regardless of OS, python version or the place where code formatting check is happening.
Installation
To start, one need to install required development dependencies (actual versions of black, flake8 and git commit hooks):
$ pip install -r requirements.dev.txt
Pre-Commit Hooks
A pre-commit hook as an easy way to ensure all files in the commit are already formatted accordingly and pass linter checks. If they are not, the git will prevent commit of problematic files unless errors are fixed.
To start, run the following command from SG repo root:
$ pip install pre-commit
$ pre-commit installThe command should complete without errors. Once done, all your upcoming commits will be checked via black & flake8.
Usage
Just run ``$ black .` from the SG root. It will reformat the whole repo.`
For flake8: $ flake8 --statistics --config scripts/flake8-config setup.py .`
Signed Commits
Background
Signed commits provide a way to verify the authenticity and integrity of the code changes made by a particular developer, as the commit is cryptographically signed using their private GPG key. This helps ensure that the code changes were made by the intended person and have not been tampered with during transit.
You can find more information here.
Add GPG key to GitHub
1. Generate a new GPG key. After completing the steps 1-9 you'll see the console message containing generated key ID (e.g. <i>3AA5C34371567BD2</i>).
2. Copy the GPG key by running the command on step 12 from the link above
$ gpg --armor --export <YOUR_GPG_KEY_ID>3. Add the new GPG key to your GitHub account
Use GPG key
- From Pycharm
- From Terminal, but first also do:
$ git config --global user.signingkey <YOUR_GPG_KEY_ID>
$ git config --global gpg.program $(which gpg)
$ git config --global commit.gpgsign trueMaking signed commits
Use `
-S` flag to sign your commit:
bashgit commit -S -m "Making my first signed commit"
You could run the next command to ensure the commit has been signed:
bashgit log --show-signature -1
I have created PR with unsigned commits. Do I have to start over?
No. What you need to do is create another branch starting from master, then move your commits from PR branch to the
new branch and force push it to the remote under the old name. Let's say your PR branch named 'feature/my_awesome_pr'
then you need to do the following:
bashgit checkout master
git checkout -b feature/my_awesome_pr_signed
git merge --squash --no-commit feature/my_awesome_pr
git commit -S -m "My signed commit"
git push -f origin feature/my_awesome_pr_signed:feature/my_awesome_pr
The last command will overwrite your PR branch with the new signed commit containing all changes from the PR.
GPG debug
Some of contributors can face this problem while making their first commit:
texterror: gpg failed to sign the data
fatal: failed to write commit object
You could consider adding `GIT_TRACE=1` at the beginning of the `git commit` command, this will show you what the commit command does under the hood. In case you are seeing something like this
bash18:57:12.099725 run-command.c:663 trace: run_command: /usr/bin/gpg --status-fd=2 -bsau <YOUR_GPG_KEY_ID>
extract the GPG command (`/usr/bin/gpg --status-fd=2 -bsau <YOUR_GPG_KEY_ID>`) and execute it separately. Now you can check what happened while running the GPG signature.Jupyter Notebooks Contribution
Pulling updates from remote might cause merge conflicts with jupyter notebooks. The tool nbdime might solve this.
* Installing nbdime
textpip install nbdime
* Run a diff between two notebooks
textnbdiff notebook_1.ipynb notebook_2.ipynb
---
LICENSE (LICENSE.md)
Apache License
Version 2.0, January 2004
http://www.apache.org/licenses/
TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
1. Definitions.
"License" shall mean the terms and conditions for use, reproduction,
and distribution as defined by Sections 1 through 9 of this document.
"Licensor" shall mean the copyright owner or entity authorized by
the copyright owner that is granting the License.
"Legal Entity" shall mean the union of the acting entity and all
other entities that control, are controlled by, or are under common
control with that entity. For the purposes of this definition,
"control" means (i) the power, direct or indirect, to cause the
direction or management of such entity, whether by contract or
otherwise, or (ii) ownership of fifty percent (50%) or more of the
outstanding shares, or (iii) beneficial ownership of such entity.
"You" (or "Your") shall mean an individual or Legal Entity
exercising permissions granted by this License.
"Source" form shall mean the preferred form for making modifications,
including but not limited to software source code, documentation
source, and configuration files.
"Object" form shall mean any form resulting from mechanical
transformation or translation of a Source form, including but
not limited to compiled object code, generated documentation,
and conversions to other media types.
"Work" shall mean the work of authorship, whether in Source or
Object form, made available under the License, as indicated by a
copyright notice that is included in or attached to the work
(an example is provided in the Appendix below).
"Derivative Works" shall mean any work, whether in Source or Object
form, that is based on (or derived from) the Work and for which the
editorial revisions, annotations, elaborations, or other modifications
represent, as a whole, an original work of authorship. For the purposes
of this License, Derivative Works shall not include works that remain
separable from, or merely link (or bind by name) to the interfaces of,
the Work and Derivative Works thereof.
"Contribution" shall mean any work of authorship, including
the original version of the Work and any modifications or additions
to that Work or Derivative Works thereof, that is intentionally
submitted to Licensor for inclusion in the Work by the copyright owner
or by an individual or Legal Entity authorized to submit on behalf of
the copyright owner. For the purposes of this definition, "submitted"
means any form of electronic, verbal, or written communication sent
to the Licensor or its representatives, including but not limited to
communication on electronic mailing lists, source code control systems,
and issue tracking systems that are managed by, or on behalf of, the
Licensor for the purpose of discussing and improving the Work, but
excluding communication that is conspicuously marked or otherwise
designated in writing by the copyright owner as "Not a Contribution."
"Contributor" shall mean Licensor and any individual or Legal Entity
on behalf of whom a Contribution has been received by Licensor and
subsequently incorporated within the Work.
2. Grant of Copyright License. Subject to the terms and conditions of
this License, each Contributor hereby grants to You a perpetual,
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
copyright license to reproduce, prepare Derivative Works of,
publicly display, publicly perform, sublicense, and distribute the
Work and such Derivative Works in Source or Object form.
3. Grant of Patent License. Subject to the terms and conditions of
this License, each Contributor hereby grants to You a perpetual,
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
(except as stated in this section) patent license to make, have made,
use, offer to sell, sell, import, and otherwise transfer the Work,
where such license applies only to those patent claims licensable
by such Contributor that are necessarily infringed by their
Contribution(s) alone or by combination of their Contribution(s)
with the Work to which such Contribution(s) was submitted. If You
institute patent litigation against any entity (including a
cross-claim or counterclaim in a lawsuit) alleging that the Work
or a Contribution incorporated within the Work constitutes direct
or contributory patent infringement, then any patent licenses
granted to You under this License for that Work shall terminate
as of the date such litigation is filed.
4. Redistribution. You may reproduce and distribute copies of the
Work or Derivative Works thereof in any medium, with or without
modifications, and in Source or Object form, provided that You
meet the following conditions:
(a) You must give any other recipients of the Work or
Derivative Works a copy of this License; and
(b) You must cause any modified files to carry prominent notices
stating that You changed the files; and
(c) You must retain, in the Source form of any Derivative Works
that You distribute, all copyright, patent, trademark, and
attribution notices from the Source form of the Work,
excluding those notices that do not pertain to any part of
the Derivative Works; and
(d) If the Work includes a "NOTICE" text file as part of its
distribution, then any Derivative Works that You distribute must
include a readable copy of the attribution notices contained
within such NOTICE file, excluding those notices that do not
pertain to any part of the Derivative Works, in at least one
of the following places: within a NOTICE text file distributed
as part of the Derivative Works; within the Source form or
documentation, if provided along with the Derivative Works; or,
within a display generated by the Derivative Works, if and
wherever such third-party notices normally appear. The contents
of the NOTICE file are for informational purposes only and
do not modify the License. You may add Your own attribution
notices within Derivative Works that You distribute, alongside
or as an addendum to the NOTICE text from the Work, provided
that such additional attribution notices cannot be construed
as modifying the License.
You may add Your own copyright statement to Your modifications and
may provide additional or different license terms and conditions
for use, reproduction, or distribution of Your modifications, or
for any such Derivative Works as a whole, provided Your use,
reproduction, and distribution of the Work otherwise complies with
the conditions stated in this License.
5. Submission of Contributions. Unless You explicitly state otherwise,
any Contribution intentionally submitted for inclusion in the Work
by You to the Licensor shall be under the terms and conditions of
this License, without any additional terms or conditions.
Notwithstanding the above, nothing herein shall supersede or modify
the terms of any separate license agreement you may have executed
with Licensor regarding such Contributions.
6. Trademarks. This License does not grant permission to use the trade
names, trademarks, service marks, or product names of the Licensor,
except as required for reasonable and customary use in describing the
origin of the Work and reproducing the content of the NOTICE file.
7. Disclaimer of Warranty. Unless required by applicable law or
agreed to in writing, Licensor provides the Work (and each
Contributor provides its Contributions) on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
implied, including, without limitation, any warranties or conditions
of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
PARTICULAR PURPOSE. You are solely responsible for determining the
appropriateness of using or redistributing the Work and assume any
risks associated with Your exercise of permissions under this License.
8. Limitation of Liability. In no event and under no legal theory,
whether in tort (including negligence), contract, or otherwise,
unless required by applicable law (such as deliberate and grossly
negligent acts) or agreed to in writing, shall any Contributor be
liable to You for damages, including any direct, indirect, special,
incidental, or consequential damages of any character arising as a
result of this License or out of the use or inability to use the
Work (including but not limited to damages for loss of goodwill,
work stoppage, computer failure or malfunction, or any and all
other commercial damages or losses), even if such Contributor
has been advised of the possibility of such damages.
9. Accepting Warranty or Additional Liability. While redistributing
the Work or Derivative Works thereof, You may choose to offer,
and charge a fee for, acceptance of support, warranty, indemnity,
or other liability obligations and/or rights consistent with this
License. However, in accepting such obligations, You may act only
on Your own behalf and on Your sole responsibility, not on behalf
of any other Contributor, and only if You agree to indemnify,
defend, and hold each Contributor harmless for any liability
incurred by, or claims asserted against, such Contributor by reason
of your accepting any such warranty or additional liability.
END OF TERMS AND CONDITIONS
APPENDIX: How to apply the Apache License to your work.
To apply the Apache License to your work, attach the following
boilerplate notice, with the fields enclosed by brackets "[]"
replaced with your own identifying information. (Don't include
the brackets!) The text should be enclosed in the appropriate
comment syntax for the file format. We also recommend that a
file or class name and description of purpose be included on the
same "printed page" as the copyright notice for easier
identification within third-party archives.
Copyright [2022] [Deci-AI]
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.
---
LICENSE.YOLONAS POSE (LICENSE.YOLONAS-POSE.md)
YOLO-NAS-POSE License
These model weights or any components comprising the model and the associated documentation (the "Software") is licensed to you by Deci.AI, Inc. ("Deci") under the following terms:
Β© 2023 β Deci.AI, Inc.
Subject to your full compliance with all of the terms herein, Deci hereby grants you a non-exclusive, revocable, non-sublicensable, non-transferable worldwide and limited right and license to use the Software. If you are using the Deci platform for model optimization, your use of the Software is subject to the Terms of Use available here (the "Terms of Use").
You shall not, without Deci's prior written consent:
(i) resell, lease, sublicense or distribute the Software to any person;
(ii) use the Software to provide third parties with managed services or provide remote access to the Software to any person or compete with Deci in any way;
(iii) represent that you possess any proprietary interest in the Software;
(iv) directly or indirectly, take any action to contest Deci's intellectual property rights or infringe them in any way;
(V) reverse-engineer, decompile, disassemble, alter, enhance, improve, add to, delete from, or otherwise modify, or derive (or attempt to derive) the technology or source code underlying any part of the Software;
(vi) use the Software (or any part thereof) in any illegal, indecent, misleading, harmful, abusive, harassing and/or disparaging manner or for any such purposes. Except as provided under the terms of any separate agreement between you and Deci, including the Terms of Use to the extent applicable, you may not use the Software for any commercial use, including in connection with any models used in a production environment.
DECI PROVIDES THE SOFTWARE "AS IS" WITHOUT WARRANTY OF ANY KIND, EITHER EXPRESSED OR IMPLIED, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE OR NON-INFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS OF THE SOFTWARE BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
---
LICENSE.YOLONAS (LICENSE.YOLONAS.md)
YOLO-NAS License
These model weights or any components comprising the model and the associated documentation (the "Software") is licensed to you by Deci.AI, Inc. ("Deci") under the following terms:
Β© 2023 β Deci.AI, Inc.
Subject to your full compliance with all of the terms herein, Deci hereby grants you a non-exclusive, revocable, non-sublicensable, non-transferable worldwide and limited right and license to use the Software. If you are using the Deci platform for model optimization, your use of the Software is subject to the Terms of Use available here (the "Terms of Use").
You shall not, without Deci's prior written consent:
(i) resell, lease, sublicense or distribute the Software to any person;
(ii) use the Software to provide third parties with managed services or provide remote access to the Software to any person or compete with Deci in any way;
(iii) represent that you possess any proprietary interest in the Software;
(iv) directly or indirectly, take any action to contest Deci's intellectual property rights or infringe them in any way;
(V) reverse-engineer, decompile, disassemble, alter, enhance, improve, add to, delete from, or otherwise modify, or derive (or attempt to derive) the technology or source code underlying any part of the Software;
(vi) use the Software (or any part thereof) in any illegal, indecent, misleading, harmful, abusive, harassing and/or disparaging manner or for any such purposes. Except as provided under the terms of any separate agreement between you and Deci, including the Terms of Use to the extent applicable, you may not use the Software for any commercial use, including in connection with any models used in a production environment.
DECI PROVIDES THE SOFTWARE "AS IS" WITHOUT WARRANTY OF ANY KIND, EITHER EXPRESSED OR IMPLIED, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE OR NON-INFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS OF THE SOFTWARE BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
---
Mkdocs.Yml (mkdocs.yml)
site_name: super-gradients
docs_dir: .
nav:
- Welcome:
- Intro: ./documentation/source/welcome.md
- Installation: ./documentation/source/installation.md
- Model Zoo: ./documentation/source/model_zoo.md
- Quick Start:
- Basic: ./documentation/source/QuickstartBasicToolkit.md
- Classification: ./documentation/source/Example_Classification.md
- Object Detection: ./documentation/source/ObjectDetection.md
- Segmentation: ./documentation/source/Segmentation.md
- Pose Estimation: ./documentation/source/PoseEstimation.md
- Training an external model: ./documentation/source/Example_Training-an-external-model.md
- Pretrained Model Prediction:
- Prediction: ./documentation/source/ModelPredictions.md
- Custom training Setup: ./documentation/source/PredictionSetup.md
- Main Components:
- Models: ./documentation/source/models.md
- Dataset:
- Data: ./documentation/source/Data.md
- Computer Vision Datasets: ./src/super_gradients/training/datasets/Dataset_Setup_Instructions.md
- Dataset Adapter: ./documentation/source/dataloader_adapter.md
- Loss functions: ./documentation/source/Losses.md
- LR Assignment: ./documentation/source/LRAssignment.md
- LR schedulers: ./documentation/source/LRScheduling.md
- Metrics: ./documentation/source/Metrics.md
- Optimizers: ./documentation/source/optimizers.md
- Phase Callbacks: ./documentation/source/PhaseCallbacks.md
- YAMLs and Recipes:
- Configurations: ./documentation/source/configuration_files.md
- Training: ./documentation/source/Recipes_Training.md
- Factories: ./documentation/source/Recipes_Factories.md
- Custom Recipes: ./documentation/source/Recipes_Custom.md
- Experiment Management: ./documentation/source/experiment_management.md
- Checkpoints: ./documentation/source/Checkpoints.md
- Docker: ./documentation/source/SGDocker.md
- Output Adapter: ./documentation/source/DetectionOutputAdapter.md
- Features:
- Training Modes: ./documentation/source/device.md
- Logging: ./documentation/source/logs.md
- Experiment Monitoring: ./documentation/source/experiment_monitoring.md
- Exponential Moving Average (EMA): ./documentation/source/EMA.md
- Automatic Mixed Precision (AMP): ./documentation/source/average_mixed_precision.md
- Knowledge Distillation (KD): ./documentation/source/KD.md
- Quantization (PTQ & QAT): ./documentation/source/ptq_qat.md
- Model Export (ONNX & TensorRT):
- Object Detection: ./documentation/source/models_export.md
- Pose Estimation: ./documentation/source/models_export_pose.md
- Troubleshooting: ./documentation/source/troubleshooting.md
- Contribution: ./documentation/source/CONTRIBUTING.md
- Code: this_is_automatically_generated.md
---
Requirements.Dev (requirements.dev.txt)
flake8==5.0.4
black==22.10.0
pre-commit==2.20.0
gitpython>=3.1.0
ipykernel==6.25
nbconvert==7.8.0
pycocotools==2.0.6
sphinx~=4.0.2
sphinx-rtd-theme
coverage~=5.3.1
---
Requirements (requirements.txt)
torch>=1.9.0
tqdm>=4.57.0
boto3>=1.17.15
jsonschema>=3.2.0
Deprecated>=1.2.11
scipy>=1.6.1
matplotlib>=3.3.4
psutil>=5.8.0
tensorboard>=2.4.1
pinned because of a crash in Snyk
setuptools>=65.5.1,<67.0.0
torchvision>=0.10.0
torchmetrics==0.8
hydra-core>=1.2.0
onnxruntime>=1.15.0
onnx 1.16.0 introduce IR 10 which is not yet supported by onnx runtime & graphsurgeon
onnx==1.15.0
pillow>=10.2.0
pip-tools>=6.12.1
einops==0.3.2
treelib==1.6.1
termcolor==1.1.0
packaging>=20.4
not directly required, pinned by Snyk to avoid a vulnerability
wheel>=0.38.0
not directly required, pinned by Snyk to avoid a vulnerability
pygments>=2.7.4
stringcase>=1.2.0
rapidfuzz
json-tricks==3.16.1
onnxsim>=0.4.3,<1.0
data-gradients~=0.3.1
albumentations~=1.3
not directly required, pinned by Snyk to avoid a vulnerability
fonttools>=4.43.0
not directly required, pinned by Snyk to avoid a vulnerability
werkzeug>=2.3.8
imagesize~=1.4.1---
YOLONAS POSE (YOLONAS-POSE.md)
YOLO-NAS-POSE
A Next-Generation, Pose Estimation Foundational Model generated by Deciβs Neural Architecture Search Technology
Deci is thrilled to announce the release of a new object detection model, YOLO-NAS-POSE - a derivative of YOLO-NAS,
pose estimation architecture, providing superior real-time object detection capabilities and production-ready performance.
Deci's mission is to provide AI teams with tools to remove development barriers and attain efficient inference performance more quickly.
The new YOLO-NAS-POSE delivers state-of-the-art (SOTA) performance with the unparalleled accuracy-speed performance, outperforming other models such as YOLOv8-Pose, DEKR and others.
Deci's proprietary Neural Architecture Search technology, AutoNACβ’, generated the architecture of YOLO-NAS-POSE model.
The AutoNACβ’ engine lets you input any task, data characteristics (access to data is not required), inference environment and performance targets,
and then guides you to find the optimal architecture that delivers the best balance between accuracy and inference speed for your specific application.
In addition to being data and hardware aware, the AutoNAC engine considers other components in the inference stack, including compilers and quantization.
| Model | AP | Latency (ms) |
|------------------|-------|--------------|
| YOLO-NAS N | 59.68 | 2.35 ms |
| YOLO-NAS S | 64.15 | 3.29 ms |
| YOLO-NAS M | 67.87 | 6.87 ms |
| YOLO-NAS L | 68.24 | 8.86 ms |
AP numbers in table reported for COCO 2017 Val dataset and latency benchmarked for 640x640 images on Nvidia T4 GPU.
No flip-TTA was used.
Similarly to YOLO-NAS, YOLO-NAS-POSE architecture employs quantization-aware blocks and selective quantization for optimized performance.
In fact YOLO-NAS-POSE is a derivative of YOLO-NAS and uses same backbone and neck as YOLO-NAS.
Only the head is different and is optimized by AutoNAC for pose estimation task.
That enables us to use transfer learning and fine-tune YOLO-NAS-POSE starting from YOLO-NAS weights.
Quickstart
Extract predicted poses
pythonimport super_gradientsyolo_nas = super_gradients.training.models.get("yolo_nas_pose_l", pretrained_weights="coco_pose").cuda()
model_predictions = yolo_nas.predict("https://deci-pretrained-models.s3.amazonaws.com/sample_images/beatles-abbeyroad.jpg", conf=0.5).show()
prediction = model_predictions[0].prediction # One prediction per image - Here we work with 1 image, so we get the first.
bboxes = prediction.bboxes_xyxy # [Num Instances, 4] List of predicted bounding boxes for each object
poses = prediction.poses # [Num Instances, Num Joints, 3] list of predicted joints for each detected object (x,y, confidence)
scores = prediction.scores # [Num Instances] - Confidence value for each predicted instance
Recipes
We provide training recipies for training YOLO-NAS-POSE on COCO, CrowdPose and AnimalPose datasets.
#### COCO 2017
* super_gradients/recipes/coco2017_yolo_nas_pose_n.yaml
* super_gradients/recipes/coco2017_yolo_nas_pose_s.yaml
* super_gradients/recipes/coco2017_yolo_nas_pose_m.yaml
* super_gradients/recipes/coco2017_yolo_nas_pose_l.yaml
Additional resources
<table>
<tr>
<td>
<a target="_blank" href="https://colab.research.google.com/drive/1O4N5Vbzv0rfkT81LQidPktX8RtoS5A40">
<img src="./documentation/assets/SG_img/colab_logo.png" /> Predict poses with YoloNAS Pose Model
</a>
</td>
</tr>
<tr>
<td>
<a href="https://colab.research.google.com/github/Deci-AI/super-gradients/blob/master/notebooks/YoloNAS_Pose_Fine_Tuning_Animals_Pose_Dataset.ipynb">
<img src="./documentation/assets/SG_img/colab_logo.png" alt="Open In Colab" />
</a>
<a href="https://github.com/Deci-AI/super-gradients/blob/master/notebooks/YoloNAS_Pose_Fine_Tuning_Animals_Pose_Dataset.ipynb">
Fine Tuning YoloNAS-Pose on AnimalPose dataset
</a>
</td>
</tr>
<tr>
<td>
<a target="_blank" href="documentation/source/YoloNASPoseQuickstart.md">
Documentation: YOLO-NAS-POSE Quickstart
</a>
</td>
</tr>
<tr>
<td>
<a target="_blank" href="documentation/source/Recipes_Training.md">
Documentation: Recipies
</a>
</td>
</tr>
<tr>
<td>
<a target="_blank" href="documentation/source/models_export_pose.md">
Documentation: YOLO-NAS-POSE Export
</a>
</td>
</tr>
<tr>
<td>
Join our <a target="_blank" href="https://discord.gg/2v6cEGMREN">
Discord Community
</a>
</td>
</tr>
</table>
LICENSE
The YOLO-NAS-POSE model is available under an open-source license with pre-trained weights available for non-commercial use on SuperGradients, Deci's PyTorch-based, open-source, computer vision training library.
With SuperGradients, users can train models from scratch or fine-tune existing ones, leveraging advanced built-in training techniques like Distributed Data Parallel, Exponential Moving Average, Automatic mixed precision, and Quantization Aware Training.
License file is available here: YOLO-NAS-POSE WEIGHTS LICENSE
---
YOLONAS (YOLONAS.md)
YOLO-NAS
A Next-Generation, Object Detection Foundational Model generated by Deciβs Neural Architecture Search Technology
Deci is thrilled to announce the release of a new object detection model, YOLO-NAS - a game-changer in the world of object detection, providing superior real-time object detection capabilities and production-ready performance. Deci's mission is to provide AI teams with tools to remove development barriers and attain efficient inference performance more quickly.
The new YOLO-NAS delivers state-of-the-art (SOTA) performance with the unparalleled accuracy-speed performance, outperforming other models such as YOLOv5, YOLOv6, YOLOv7 and YOLOv8.
Deci's proprietary Neural Architecture Search technology, AutoNACβ’, generated the YOLO-NAS model. The AutoNACβ’ engine lets you input any task, data characteristics (access to data is not required), inference environment and performance targets, and then guides you to find the optimal architecture that delivers the best balance between accuracy and inference speed for your specific application. In addition to being data and hardware aware, the AutoNAC engine considers other components in the inference stack, including compilers and quantization.
In terms of pure numbers, YOLO-NAS is ~0.5 mAP point more accurate and 10-20% faster than equivalent variants of YOLOv8 and YOLOv7.
| Model | mAP | Latency (ms) |
|------------------|-------|--------------|
| YOLO-NAS S | 47.5 | 3.21 |
| YOLO-NAS M | 51.55 | 5.85 |
| YOLO-NAS L | 52.22 | 7.87 |
| YOLO-NAS S INT-8 | 47.03 | 2.36 |
| YOLO-NAS M INT-8 | 51.0 | 3.78 |
| YOLO-NAS L INT-8 | 52.1 | 4.78 |
mAP numbers in table reported for Coco 2017 Val dataset and latency benchmarked for 640x640 images on Nvidia T4 GPU.
YOLO-NAS's architecture employs quantization-aware blocks and selective quantization for optimized performance. When converted to its INT8 quantized version, YOLO-NAS experiences a smaller precision drop (0.51, 0.65, and 0.45 points of mAP for S, M, and L variants) compared to other models that lose 1-2 mAP points during quantization. These techniques culminate in innovative architecture with superior object detection capabilities and top-notch performance.
Quickstart
Extract bounding boxes
pythonimport super_gradientsyolo_nas = super_gradients.training.models.get("yolo_nas_l", pretrained_weights="coco").cuda()
model_predictions = yolo_nas.predict("https://deci-pretrained-models.s3.amazonaws.com/sample_images/beatles-abbeyroad.jpg").show()
prediction = model_predictions[0].prediction # One prediction per image - Here we work with 1 image so we get the first.
bboxes = prediction.bboxes_xyxy # [[Xmin,Ymin,Xmax,Ymax],..] list of all annotation(s) for detected object(s)
bboxes = prediction.bboxes_xyxy # [[Xmin,Ymin,Xmax,Ymax],..] list of all annotation(s) for detected object(s)
class_names = prediction.class_names # ['Class1', 'Class2', ...] List of the class names
class_name_indexes = prediction.labels.astype(int) # [2, 3, 1, 1, 2, ....] Index of each detected object in class_names(corresponding to each bounding box)
confidences = prediction.confidence.astype(float) # [0.3, 0.1, 0.9, ...] Confidence value(s) in float for each bounding boxes
Recipes
We provide fine-tuning recipies for Roboflow-100 datasets.
* super_gradients/recipes/roboflow_yolo_nas_m.yaml
* super_gradients/recipes/roboflow_yolo_nas_s.yaml
* super_gradients/recipes/roboflow_yolo_nas_s_qat.yaml
Great fine-tuning potential
We demonstrate great performance of YOLO-NAS on downstream tasks. When fine-tuning on Roboflow-100 our YOLO-NAS model achieves higher mAP than our nearest competitors:
Additional resources
<table>
<tr>
<td>
<a target="_blank" href="https://bit.ly/yolo-nas-starter-notebook">
<img src="./documentation/assets/SG_img/colab_logo.png" /> Fine-Tuning Notebook
</a>
</td>
</tr><tr>
<td>
<a target="_blank" href="https://bit.ly/3MIKdTy">
<img src="./documentation/assets/SG_img/colab_logo.png" /> Quantization Aware Training YoloNAS on Custom Dataset Notebook
</a>
</td>
</tr>
<tr>
<td>
<a target="_blank" href="documentation/source/YoloNASQuickstart.md">
Documentation: YOLO-NAS Quickstart
</a>
</td>
</tr>
<tr>
<td>
<a target="_blank" href="documentation/source/qat_ptq_yolo_nas.md">
Documentation: YOLO-NAS Quantization-Aware training and post-training Quantization
</a>
</td>
</tr><tr>
<td>
Join our <a target="_blank" href="https://discord.gg/2v6cEGMREN">
Discord Community
</a>
</td>
</tr>
</table>
LICENSE
The YOLO-NAS model is available under an open-source license with pre-trained weights available for non-commercial use on SuperGradients, Deci's PyTorch-based, open-source, computer vision training library.
With SuperGradients, users can train models from scratch or fine-tune existing ones, leveraging advanced built-in training techniques like Distributed Data Parallel, Exponential Moving Average, Automatic mixed precision, and Quantization Aware Training.
License file is available here: YOLO-NAS WEIGHTS LICENSE
---
.Pre Commit Config.Yaml (.pre-commit-config.yaml)
repos:
- repo: https://github.com/pre-commit/pre-commit-hooks
rev: v2.3.0
hooks:
- id: check-yaml
- id: check-toml
- id: end-of-file-fixer
- repo: https://github.com/psf/black
rev: 22.10.0
hooks:
- id: black
args: [ --config=pyproject.toml ]
- repo: https://github.com/pycqa/flake8
rev: 5.0.4
hooks:
- id: flake8
args: [ --config=scripts/flake8-config ]
---
Documentation/Source/Average Mixed Precision (documentation/source/average_mixed_precision.md)
Automatic Mixed Precision (AMP)
Automatic mixed precision (AMP) is a feature in PyTorch that enables the use of lower-precision data types, such as float16, in deep learning models for improved memory and computation efficiency.
It automatically casts the model's parameters and buffers to a lower-precision data type, and dynamically rescales the activations to prevent underflow or overflow.
Set up AMP
To use AMP in SuperGradients, you simply need to set mixed_precision=True in your training_params.In python script
pythonfrom super_gradients import Trainertrainer = Trainer("experiment_name")
model = ...
training_params = {"mixed_precision": True, ...:...}
trainer.train(model=model, training_params=training_params, ...)
In recipe
yamlmy_training_hyperparams.yaml
mixed_precision: True # Whether to use mixed precision or not.
---
Documentation/Source/BenchmarkingYoloNAS (documentation/source/BenchmarkingYoloNAS.md)
Benchmarking YoloNAS
TLDR: Getting 10FPS for YoloNAS on your 4090 and feeling cheated? Read this carefully!
Introduction:
YoloNAS is a leading object detection architecture that combines accuracy and efficiency. Using post training quantization (PTQ) and quantization-aware training (QAT) YoloNAS models can be optimized for resource-constrained devices.
However, to fully tap into its potential, it is crucial to know how to export the quantized model to the INT8 TensorRT (TRT) engine.
In this tutorial, we emphasize the significance of this step and provide a concise guide to efficiently exporting a quantized YoloNAS model to the INT8 TRT engine.
Doing so teaches us how to properly benchmark YoloNAS and understand its full potential.
Step 1: Export YoloNAS to ONNX
The first step is to export our YoloNAS model to ONNX correctly. Two actions must be taken before we export our model to onnx:
1. We must call
model.prep_model_for_conversion - this is essential as YoloNAS incorporates QARepVGG blocks. Without this call, the RepVGG branches will not be fused, and our model's speed will decrease significantly! This is true for the Pytorch model as well as the compiled TRT Engine!
Nothing to worry about if you have quantized your model with PTQ/QAT with SG, as this is done under the hood before exporting the ONNX checkpoints.
2. We need to replace our layers with "fake quantized" ones - this happens when we perform post-training quantization or quantization-aware training with SG.
Again, nothing to worry about if you performed PTQ/QAT with SG and hold your newly exported ONNX checkpoint. Beware that inference time in Pytorch is slower with such blocks - but will be faster once converted to the TRT Engine.
There are plenty of guides on how to perform PTQ/QAT with SG:
- Quantization-aware fine-tuning YoloNAS on custom dataset notebook
- QA/PTQ YoloNAS with configuration files
- QA/PTQ
Suppose we ran PTQ/QAT, then our PTQ/QAT checkpoints have been exported to our checkpoints directory.
If we plug them into netron, we can see that new blocks that were not a part of the original network were introduced: the Quantize/Dequantize layers -
<div>
<img src="images/qdq_yolonas_netron.png" width="750">
</div>
This is expected and an excellent way to verify that our model is ready to be converted to Int8 using Nvidia's TesnorRT.
As stated earlier - inference time in Pytorch is slower with such blocks - but will be faster once converted to the TRT Engine.
Step 2: Create TRT Engine
First, please make sure to install Nvidia's TensorRT.
TensorRT version >= 8.4 is required.
We can now use these ONNX files to deploy our newly trained YoloNAS models to production. When building the TRT engine, it is essential to specify that we convert to Int8 (the fake quantized layers in our models will be adapted accordingly); this can be done by running:
commandlinetrtexec --fp16 --int8 --avgRuns=100 --onnx=your_yolonas_qat_model.onnx
Step 3: View Model Benchmark Results
After running:
commandlinetrtexec --fp16 --int8 --avgRuns=100 --onnx=your_yolonas_qat_model.onnx
your screen will look somewhat similar to the screenshot below:
<div>
<img src="images/trtexec.png" width="750">
</div>Command notes:
- Note that this process might take some time, depending on the GPU, the model, and the size of the input (up to 40 minutes is reasonable on smaller devices).
- Also notice the
--avgRuns=100 which means that this command runs the model 100 times, so that we get more "robust" results that are less affected by noise.
- Note that since we ran PTQ/QAT we need the --int8 flag flag. But if we did not do so, then the --int8 flag will dramatically degrade the accuracy of our compiled model.
- By passing --fp16 and --int together, we allow the hybrid quantization (that is, some layers are quantized while others are not).Benchmark breakdown:
- The actual throughput and latency of your model are in blue. This tells you how your model is actually performing.
- Note that trtexec shows the minimum, maximum, mean, and median values. Significant differences between these values can indicate that your measurements are noisy. It might be that some other process uses the GPU while benchmarking or that the GPU needs to be adequately cooled.
- The end-to-end latency is marked in yellow. This includes the time it takes to prepare the input and pass it to the GPU, the GPU compute time, and the time it takes to move the output from the GPU back to the host (for the full batch size). If you plan on running batches one by one synchronously, this is the time that affects you. But if you use an async inference engine, you will be affected by the numbers in blue.
- High H2D (H=Host=CPU; D=Device=GPU) values indicate your input size has a crucial effect on the performance. Consider resizing the input in advance (not on the GPU), or test with different batch sizes to find the optimal setting.
- High D2H values indicate your output might be too big. You can consider a task-specific method to reduce it. i.e., use top-k at the end of your detection model to limit the number of boxes coming out. Alternatively, use a softmax layer at the end of your segmentation model to change the output representation to one with smaller dimensions.
---
Documentation/Source/Checkpoints (documentation/source/Checkpoints.md)
Model Checkpoints
If you are not familiar on how experiments are managed, you can check this tutorial
The first question that arises is: what is a checkpoint?
From the Pytorch Lightning documentation:
When a model is training, the performance changes as it sees more data. It is a best practice to save the state of a model throughout the training process. This gives you a version of the model, a checkpoint, at each key point during the development of the model. Once training has been completed, use the checkpoint corresponding to the best performance you found during the training process.
Checkpoints also enable your training to resume where it was in case the training process is interrupted.
Checkpoints Saving: Which, When, and Where?
From the previous subsection, we understand that checkpoints saved at different times during training have different purposes.
That's why in SG, multiple checkpoints are saved throughout training:
| Checkpoint Filename | When is it saved?|
| ------------- |:-------------:|
|
ckpt_best.pth | Each time we reach a new best metric_to_watch when performing validation. |
| ckpt_latest.pth | At the end of every epoch, constantly overriding. |
| average_model.pth | At the end of training - composed of 10 best models according to metric_to_watch and will only be save when the training_param average_best_models=True. |
| ckpt_epoch_{EPOCH_INDEX}.pth | At the end of a fixed epoch number EPOCH_INDEX if it is specified through save_ckpt_epoch_list training_param |#### Where are the checkpoint files saved?
The checkpoint files will be saved at
<ckpt_root_dir>/<experiment_name>/<run_dir>.-
ckpt_root_dir and experiment_name can be set by the user when instantiating the Trainer.
pythonTrainer(ckpt_root_dir='path/to/ckpt_root_dir', experiment_name="my_experiment")
- run_dir is unique and automatically generated each time you start a new training, with trainer.train(...)When working with a cloned version of SG, one can leave out the
ckpt_root_dir arg, and checkpoints will be saved under the super_gradients/checkpoints directory.Checkpoint Structure
Checkpoints in SG are instances of state_dict.
They hold additional information about the model's training besides the model weights.
The checkpoint keys:
-
net: The network's state_dict (state_dict).
- acc: The network's achieved metric value on the validation set (metric_to_watch in training_params (float).
- epoch: The last epoch performed.
- optimizer_state_dict: The state_dict of the optimizer (state_dict).
- scaler_state_dict: Optional - only present when training with mixed_precision=True. The state_dict of Trainer.scaler.
- ema_net: Optional - only present when training with ema=True. The EMA model's state_dict. Note that average_model.pth lacks this entry even if ema=True since the average model's snapshots are of the EMA network already (i.e., the "net" entry is already an average of the EMA snapshots).
- torch_scheduler_state_dict: Optional, will only be present when using a torch native lr scheduler (see LRScheduling)
Remote Checkpoint Saving with SG Loggers
SG supports remote checkpoint saving using 3rd party tools (for example, Weights & Biases).
To do so, specify
save_checkpoints_remote=True inside sg_logger_params training_param.
See our documentation on Third-party experiment monitoring.
Loading Checkpoints
When discussing checkpoint loading in SG, we must separate it into two use cases: loading weights and resuming training.
While the former requires the model's state_dict alone, SG checkpoint loading methods introduce more functionality than PyTorch's vanilla
load_state_dict(), especially for SG-trained checkpoints.Loading Model Weights from a Checkpoint
Loading model weights can be done right after model initialization, using
models.get(...), or by explicitly calling load_checkpoint_to_model on the torch.nn.Module instance.Suppose we have launched a training experiment with a similar structure to the one below:
pythonfrom super_gradients.training import Trainer
...
...
from super_gradients.training import models
from super_gradients.common.object_names import Modelstrainer = Trainer("my_resnet18_training_experiment", ckpt_root_dir="/path/to/my_checkpoints_folder")
train_dataloader = ...
valid_dataloader = ...
model = models.get(model_name=Models.RESNET18, num_classes=10)
train_params = {
...
"loss": "CrossEntropyLoss",
"criterion_params": {},
"save_ckpt_epoch_list": [10,15]
...
}
trainer.train(model=model, training_params=train_params, train_loader=train_dataloader, valid_loader=valid_dataloader)
Then at the end of the training, our
ckpt_root_dir contents will look similar to the following:text<ckpt_root_dir>
β
βββ <experiment_name>
β β
β ββββ <run_dir>
β β ββ ckpt_best.pth # Best performance during validation
β β ββ ckpt_latest.pth # End of the most recent epoch
β β ββ average_model.pth # Averaged over specified epochs
β β ββ ckpt_epoch_*.pth # Checkpoints from specific epochs (like epoch 10, 15, etc.)
β β ββ events.out.tfevents.* # Tensorflow run artifacts
β β ββ log_<timestamp>.txt # Trainer logs of the specific run
β β
β ββββ <other_run_dir>
β ββ ...
β
ββββ <other_experiment_name>
β
ββββ <run_dir>
β ββ ...
β
ββββ <another_run_dir>
ββ ...
Suppose we wish to load the weights from
ckpt_best.pth. We can simply pass its path to the checkpoint_path argument in models.get(...):pythonfrom super_gradients.training import models
from super_gradients.common.object_names import Modelsmodel = models.get(model_name=Models.RESNET18, num_classes=10, checkpoint_path="/path/to/my_checkpoints_folder/my_resnet18_training_experiment/RUN_20230802_131052_651906/ckpt_best.pth")
Important: when loading SG-trained checkpoints using models.get(...), if the network was trained with EMA, the EMA weights will be the ones loaded.
If we already possess an instance of our model, we can also directly use
load_checkpoint_to_model:pythonfrom super_gradients.training import models
from super_gradients.common.object_names import Models
from super_gradients.training.utils.checkpoint_utils import load_checkpoint_to_modelmodel = models.get(model_name=Models.RESNET18, num_classes=10)
load_checkpoint_to_model(net=model, ckpt_local_path="/path/to/my_checkpoints_folder/my_resnet18_training_experiment/RUN_20230802_131052_651906/ckpt_best.pth")
Extending the Functionality of PyTorch's
strict Parameter in load_state_dict() load_state_dict()')" title="Copy section prompt for LLMs"> Copy SectionWhen not familiar with PyTorch's
strict parameter in load_state_dict(), please see PyTorch's docs on this matter first.The equivalent arguments for PyTorch's
strict parameter in load_state_dict() in models.get() and load_checkpoint_to_model are strict and strict_load respectively, and expect SG's StrictLoad enum type.Let's have a look at its possible values:
pythonclass StrictLoad(Enum):
"""
Wrapper for adding more functionality to torch's strict_load parameter in load_state_dict().
Attributes:
OFF - Native torch "strict_load = off" behavior. See nn.Module.load_state_dict() documentation for more details.
ON - Native torch "strict_load = on" behavior. See nn.Module.load_state_dict() documentation for more details.
NO_KEY_MATCHING - Allows the usage of SuperGradient's adapt_checkpoint function, which loads a checkpoint by matching each
layer's shapes (and bypasses the strict matching of the names of each layer (i.e., disregards the state_dict key matching)).
""" OFF = False
ON = True
NO_KEY_MATCHING = "no_key_matching"
In other words, we added another loading mode option-
no_key_matching. This option exploits the fact that the state_dicts are OrderedDicts, and comes in handy when the underlying network's structure remains the same, but the state_dicts keys do not match the ones inside the models state_dict.
Let's demonstrate the different strict modes with a simple example:pythonimport torchclass ModelA(torch.nn.Module):
def __init__(self):
super(ModelA, self).__init__()
self.conv1 = torch.nn.Conv2d(3, 6, 5)
self.conv2 = torch.nn.Conv2d(6, 16, 5)
class ModelB(torch.nn.Module):
def __init__(self):
super(ModelB, self).__init__()
self.conv1 = torch.nn.Conv2d(3, 6, 5)
self.CONV2 = torch.nn.Sequential([torch.nn.Conv2d(6, 16, 5)])
Notice the above networks have identical weight structures but will have different keys in their
state_dicts.
This is why loading a checkpoint from either one to the other, using strict=True, will fail and crash. Using strict=False will not crash and successfully load the first layer's weights only.
Using SG's no_key_matching will successfully load a checkpoint from either one to the other.Loading Pretrained Weights from the Model Zoo
Using
models.get(...), you can load any of our pre-trained models in 3 lines of code:
pythonfrom super_gradients.training import models
from super_gradients.common.object_names import Modelsmodel = models.get(Models.YOLOX_S, pretrained_weights="coco")
The
pretrained_weights argument specifies the dataset on which the pre-trained weights were trained. Here is the complete list of pre-trained weights.
Loading Checkpoints: Training with Configuration Files
Prerequisites: Training with Configuration Files
Recall the SGs recipes library structure:
The
super_gradients/recipes include the following subdirectories- arch_params - containing configuration files for instantiating different models
- checkpoint_params - containing configuration files that define the loaded and saved checkpoints parameters for the training
- conversion_params - containing configuration files for the model conversion scripts (for deployment)
- dataset_params - containing configuration files for instantiating different datasets and dataloaders
- training_hyperparams - containing configuration files holding hyper-parameters for specific recipes
And now, let's take a look at the default parameters in
checkpoint_params:yamlload_backbone: False # whether to load only the backbone part of the checkpoint
checkpoint_path: # checkpoint path that is located in super_gradients/checkpoints
strict_load: True # key matching strictness for loading checkpoint's weights
pretrained_weights: # a string describing the dataset of the pre-trained weights (for example, "imagenent").
And note the above parameters are used to start the training with different weights (fine-tuning etc.) - they are passed to model.get() in the underlying flow of
Trainer.train_from_config(...):
python@classmethod
def train_from_config(cls, cfg: Union[DictConfig, dict]) -> Tuple[nn.Module, Tuple]:
... # BUILD NETWORK
model = models.get(
...
strict_load=cfg.checkpoint_params.strict_load,
pretrained_weights=cfg.checkpoint_params.pretrained_weights,
checkpoint_path=cfg.checkpoint_params.checkpoint_path,
load_backbone=cfg.checkpoint_params.load_backbone,
)
# INSTANTIATE DATA LOADERS
train_dataloader = ...
val_dataloader = ...
...
# TRAIN
res = trainer.train(...)
...
Resuming Training
Resuming training in SG is a comprehensive process, controlled by three primary parameters that allow flexibility
in continuing or branching off from specific training checkpoints.
These parameters are used within
training_params: resume, run_id, and resume_path.yamlresume: False # Option to continue training from the latest checkpoint.
run_id: # ID to resume from a specific run within the same experiment.
resume_path: # Direct path to a specific checkpoint file (.pth) to resume training....
#### 1. Resuming the Latest Run
By setting
resume=True, SuperGradients will resume training from the last checkpoint within the same experiment.Example:
shellContinues from the latest run in the cifar_experiment.
python -m super_gradients.train_from_recipe --config-name=cifar10_resnet experiment_name=cifar_experiment training_hyperparams.resume=True
#### 2. Resuming a Specific Run
Using
run_id, you can resume training from a specific run within the same experiment, identified by the run ID.Example:
shellContinues from a specific run identified by the ID within cifar_experiment.
python -m super_gradients.train_from_recipe --config-name=cifar10_resnet experiment_name=cifar_experiment run_id=RUN_20230802_131052_651906
#### 3. Branching off from a specific checkpoint
By specifying a
resume_path, SuperGradients will create a new run directory,
allowing training to resume from that specific checkpoint,
and subsequently save the new checkpoints in this new directory.Example:
shellBranches from a specific checkpoint, creating a new run.
python -m super_gradients.train_from_recipe --config-name=cifar10_resnet experiment_name=cifar_experiment training_hyperparams.resume_path=/path/to/checkpoint.pth
#### 4. Resuming with original recipe
Resuming is parameter dependant - you cannot resume the training of a model if there is a mismatch between the model architecture defined in your recipe, and the one in your checkpoint.
Therefore, if you trained a model a while ago, and that in the meantime you changed the model architecture definition,
then you won't be able to resume its training, loading the model would simply raise an exception.
To avoid this issue, SuperGradients provides an option to resume a training based on the recipe that was originally used to train the model.
textTrainer.resume_experiment(ckpt_root_dir=..., experiment_name=..., run_id=...)
-
run_id is optional. You can use it to chose which run you want to resume. By default, it will resume the latest run of your experiment.Note that
Trainer.resume_experiment can only resume training that were launched with Trainer.train_from_config.See usage in our resume_experiment_example.
Resuming Training from SG Logger's Remote Storage (WandB only)
SG supports saving checkpoints throughout the training process in the remote storage defined by
SG Logger (more info about this object and it's role during training in SG at Third-party experiment monitoring.)
Suppose we run an experiment with a WandB SG logger, then our training_hyperparams should hold:
yamlsg_logger: wandb_sg_logger, # Weights&Biases Logger, see class super_gradients.common.sg_loggers.wandb_sg_logger.WandBSGLogger for details
sg_logger_params: # Params that will be passes to __init__ of the logger super_gradients.common.sg_loggers.wandb_sg_logger.WandBSGLogger
project_name: project_name, # W&B project name
save_checkpoints_remote: True,
save_tensorboard_remote: True,
save_logs_remote: True,
entity: <YOUR-ENTITY-NAME>, # username or team name where you're sending runs
api_server: <OPTIONAL-WANDB-URL> # Optional: In case your experiment tracking is not hosted at wandb servers
The
save_checkpoints_remote flag is set which will result in saving checkpoints in WandB throughout training.
Now, in case the training was interrupted, we can resume it from the checkpoint located in the WandB run storage by setting 2 training hyperparameters:
1. Set resume_from_remote_sg_logger:
yamlresume_from_remote_sg_logger: True
2. Pass run_id through wandb_id to sg_logger_params:
yamlsg_logger: wandb_sg_logger, # Weights&Biases Logger, see class super_gradients.common.sg_loggers.wandb_sg_logger.WandBSGLogger for details
sg_logger_params: # Params that will be passes to __init__ of the logger super_gradients.common.sg_loggers.wandb_sg_logger.WandBSGLogger
wandb_id: <YOUR_RUN_ID>
project_name: project_name, # W&B project name
save_checkpoints_remote: True,
save_tensorboard_remote: True,
save_logs_remote: True,
entity: <YOUR-ENTITY-NAME>, # username or team name where you're sending runs
api_server: <OPTIONAL-WANDB-URL> # Optional: In case your experiment tracking is not hosted at wandb servers
And that's it! Once you re-launch your training,
ckpt_latest.pth (by default) will be downloaded to the checkpoints directory, and the training will resume from it just as if it was locally stored.Evaluating Checkpoints
Analogically to the previous section, we often want to evaluate a checkpoint seamlessly without being familiar with the training configuration.
For this reason, SG introduces two methods:
Trainer.evaluate_checkpoint(...) and Trainer.evaluate_recipe(...) and play similar roles to the two previous ways of resuming experiments suggested in the last section:Trainer.evaluate_checkpoint is used to evaluate a checkpoint resulting from one of your previous experiments, using the same parameters (dataset, valid_metrics,...) as used during the training of the experiment.
Trainer.evaluate_recipe is used to evaluate a checkpoint from SGs pre-trained model zoo or to evaluate a checkpoint with different parameters.
See both usages and documentation in the corresponding scripts evaluate_checkpoint and evaluate_recipe.---
Documentation/Source/Configuration Files (documentation/source/configuration_files.md)
Configuration Files and Recipes
SuperGradients supports YAML formatted configuration files.
These files can contain training hyper-parameters, architecture parameters, datasets parameters and any
other parameters required by the training process.These parameters will be consumed as dictionaries or as function arguments by different parts of SuperGradients.
These YAML files act like a cookbook for training models, which is why they are called Recipes.
SuperGradients was designed to expose as many parameters as possible to allow outside configuration without writing a single line of code.
You can control the learning-rate, the weight-decay or even the loss function and metric used in the training, but moreover, you can even control
which block-type or activation function to use in your model. You can learn about and define all of these parameters from the configuration files.
Here is an example YAML file (training hyper-parameters in this case):
yamldefaults:
- default_train_paramsmax_epochs: 250
lr_updates:
_target_: numpy.arange
start: 100
stop: 250
step: 50
lr_decay_factor: 0.1
lr_mode: step
lr_warmup_epochs: 0
initial_lr: 0.1
loss: CrossEntropyLoss
optimizer: SGD
criterion_params: {}
optimizer_params:
weight_decay: 1e-4
momentum: 0.9
NOTE: You can use SuperGradients without using any configuration files, look into the examples directory to see how.
Why using configuration files
Using configuration file might seem too complicated or redundant at first. But, after a short training, you will find it extremely convenient and useful. Configuration file can help you manage your assets, such as datasets, models and training recipes. Keeping your code files as clean of parameters as possible,
allows you to have all of your configuration in one place and reuse the same code to define different objects.
In the following example, we define a training set and a validation set of Imagenet. both use the same piece of code
with different configurations:
yamltrain_dataset_params:
root: /data/Imagenet/train
transforms:
- RandomResizedCropAndInterpolation:
size: 224
interpolation: default
- RandomHorizontalFlip
- ToTensor
- Normalize:
mean: [0.485, 0.456, 0.406] # mean for normalization
std: [0.229, 0.224, 0.225] # std for normalizationval_dataset_params:
root: /data/Imagenet/val
transforms:
- Resize:
size: 256
- CenterCrop:
size: 224
- ToTensor
- Normalize:
mean: [0.485, 0.456, 0.406] # mean for normalization
std: [0.229, 0.224, 0.225] # std for normalization
Configuration file can also help you track the exact settings used for each one of your experiments, tweak and tune these settings, and share them with others.
Concentrating all of these configuration parameters in one place, gives you great visibility and control of your experiments.
How to use configuration files
So, if you got so far, we have probably manged to convince you that configuration files are awsome and powerful tools - welcome aboard!YAML is a human-readable data-serialization language. It is commonly used for configuration files and in applications where data is being
stored or transmitted (Wikipedia).
We parse each file into dictionaries, lists, and objects, and pass them to the code either as a recursive dictionary or as function arguments.
Let's try running a training session from a configuration file.
shellpython -m super_gradients.train_from_recipe --config-name=cifar10_resnet
You can stop the training after a few cycles. The recipe you have just used is a configuration file containing everything SG needs to know in order to train
Resnet18 on Cifar10. The actual YAML file is located in
src/super_gradients/recipes/cifar10_resnet.yaml.
In the same recipes library you can find many more configuration files defining different models, datasets,
and training hyper-parameters.
Hydra
Hydra is an open-source Python framework that provides us with many useful functionalities for YAML management. You can learn about Hydra
here. We use Hydra to load YAML files and convert them into dictionaries, while
instantiating the objects referenced in the YAML.
You can see this in the code:pythonimport hydra
from omegaconf import DictConfig@hydra.main(config_path="recipes", version_base="1.2")
def main(cfg: DictConfig) -> None:
print(cfg.experiment_name)
The
@hydra.main decorator is looking for YAML files in the super_gradients.recipes according to the name of the configuration file provided
in the first arg of the command line. In the experiment directory a
.hydra subdirectory will be created. The configuration files related to this run will be saved by hydra to that subdirectory. Two Hydra features worth mentioning are Command-Line Overrides
and YAML Composition.
We strongly recommend you to have a look at both of these pages.
Conclusion
This brief introduction has given you a glimpse into the functionality and importance of recipes within SuperGradients:
- Recipes Overview: Configuration files in YAML format that allow streamlined training and customization.
- SuperGradients' Utilization: Enhancing reproducibility, flexibility, and efficiency in defining models, datasets, and hyperparameters.
- Introduction to training: A simple demonstration of initiating a training session using a specific recipe.Next Step: More details await in the upcoming tutorials, where we'll explore more in-depth training from recipes,
and the customization, structure, and deeper functionality of recipes within SuperGradients.
---
Documentation/Source/Data (documentation/source/Data.md)
Data
To handle data, SuperGradients takes use of two Pytorch primitives:
torch.utils.data.Dataset - which is in charge of generating the samples and their corresponding labels,
and torch.utils.data.DataLoader - that wraps an iterable around the Dataset to enable easy access to the samples. In other words, torch.utils.data.Dataset defines how to load a single sample,
while torch.utils.data.DataLoader defines how to load batches of samples. For more information, see PyTorch documentation.Datasets
SuperGradients holds common public
torch.utils.data.Dataset implementations for various tasks: Classification:
Cifar10
Cifar100
ImageNetDataset
Object Detection:
COCODetectionDataset
DetectionDataset
PascalVOCDetectionDataset
Semantic Segmentation:
CoCoSegmentationDataSet
PascalAUG2012SegmentationDataSet
PascalVOC2012SegmentationDataSet
CityscapesDataset
SuperviselyPersonsDataset
PascalVOCAndAUGUnifiedDataset
Pose Estimation:
COCOKeypointsDataset
All of which can be imported from the
super_gradients.training.datasets module. Note that some of the above implementations require following a few simple setup steps, which are all documented hereCreating a
torch.utils.data.DataLoader from a dataset can be tricky, especially when defining some parameters on the fly. For example, in distributed training (i.e., DDP)
the torch.utils.data.DataLoader must be given a proper Sampler such that the dataset indices will be divided among the different processes.
textWarning: Using the wrong sampler when defining a data loader to be used with DDP will lead the different processes to iterate over the same data samples giving little to no speedup over single GPU training!
This is where SG's training.dataloaders.get comes in handy by taking the burden of instantiating the proper default sampler according to the training settings.
Once instantiated, any of the above can be passed to the torch.utils.data.DataLoader constructor and be used for training, validation, or testing:pythonfrom my_dataset import MyDataset
from super_gradients.training import dataloaders
import torchvision.transforms as T
from super_gradients.training import Trainer
from super_gradients.training.metrics import Accuracytrainer = Trainer("my_experiment")
train_dataset = MyDataset(split="train", transforms=T.ToTensor())
valid_dataset = MyDataset(split="validation", transforms=T.ToTensor())
test_dataset = MyDataset(split="test", transforms=T.ToTensor())
train_dataloader = dataloaders.get(dataset=train_dataset, dataloader_params={"batch_size": 4})
valid_dataloader = dataloaders.get(dataset=valid_dataset, dataloader_params={"batch_size": 16})
test_dataloader = dataloaders.get(dataset=test_dataset, dataloader_params={"batch_size": 16})
model = ...
train_params = {...}
trainer.train(model=model, training_params=train_params, train_loader=train_dataloader, valid_loader=valid_dataloader)
trainer.test(model=trainer.net, test_loader=test_dataloader, test_metrics_list=[Accuracy()])
Note that
dataloader_params will be unpacked in the torch.utils.data.DataLoader constructor after setting a proper sampler if one is not explicitly set.DataLoaders
As mentioned above, once instantiated, the
torch.utils.data.DataLoader objects form batches.
Therefore- these are the objects being passed to Trainer.train(...):python...
trainer = Trainer("my_experiment")
train_dataloader = ...
valid_dataloader = ...
model = ...
train_params = {...}trainer.train(model=model, training_params=train_params, train_loader=train_dataloader, valid_loader=valid_dataloader)
For your convenience, SuperGradients gives full access to all data loader objects used in our training recipes.
These are simply the
torch.utils.data.DataLoader configured by the recipe's dataset_params: cifar10_val
cifar10_train
cifar100_val
cifar100_train
coco2017_train
coco2017_val
coco2017_train_ssd_lite_mobilenet_v2
coco2017_val_ssd_lite_mobilenet_v2
imagenet_train
imagenet_val
imagenet_efficientnet_train
imagenet_efficientnet_val
imagenet_mobilenetv2_train
imagenet_mobilenetv2_val
imagenet_mobilenetv3_train
imagenet_mobilenetv3_val
imagenet_regnetY_train
imagenet_regnetY_val
imagenet_resnet50_train
imagenet_resnet50_val
imagenet_resnet50_kd_train
imagenet_resnet50_kd_val
imagenet_vit_base_train
imagenet_vit_base_val
tiny_imagenet_train
tiny_imagenet_val
pascal_aug_segmentation_train
pascal_aug_segmentation_val
pascal_voc_segmentation_train
pascal_voc_segmentation_val
supervisely_persons_train
supervisely_persons_val
pascal_voc_detection_train
pascal_voc_detection_val
These DataLoader can be imported from the super_gradients.training.dataloaders module.
Please note that these Dataset and DataLoader objects are already pre-defined with parameters required for specific training recipes. You can override these default parameters by passing two named arguments: dataset_params and dataloader_params(both of which are dictionaries), which will override the recipe settings. To learn which parameters you can override for each object, please refer to the YAML file with the same name.
For example, the code below will instantiate the data loader used for training in our
imagenet_resnet50 recipe
(including all data augmentations and any other data-related setting which we defined for training Resnet50 on Imagenet)
but changing the batch size for our needs.
We can then, also with a one-liner, instantiate the validation dataloader and call train() as always:pythonfrom super_gradients.training.dataloaders import imagenet_resnet50_train, imagenet_resnet50_val
from super_gradients.training import Trainertrain_dataloader = imagenet_resnet50_train(dataloader_params={"batch_size": 4, "shuffle": True}, dataset_params={"root": "/my_data_dir/Imagenet/train"})
valid_dataloader = imagenet_resnet50_val(dataloader_params={"batch_size": 16}, dataset_params={"root": "/my_data_dir/Imagenet/val"})
...
trainer = Trainer("my_imagenet_training_experiment")
model = ...
train_params = {...}
trainer.train(model=model, training_params=train_params, train_loader=train_dataloader, valid_loader=valid_dataloader)
DataLoaders - Training with Configuration Files
If you are still getting familiar with training with configuration files, follow this link.
Their names can reference any of the SG-predefined data loaders listed earlier.
For example, using the imagenet_resnet50_train and imagenet_resnet50_val:
yamldataset_params: ...
...
train_dataloader: imagenet_resnet50_train
val_dataloader: imagenet_resnet50_val...
Now, on the structure of
dataset_params:
yamltrain_dataset_params:
train_dataloader_params:
val_dataset_params:
val_dataloader_params:
As their names suggest - the parameters under train_dataset_params will be passed to the Dataset, and the parameters under train_dataloader_params will be given to the DataLoader.
As in the previous subsection, both train_dataloader_params and train_dataset_params will override the corresponding
parameters defined for the predefined data loader ( in our case, imagenet_resnet50 recipe's dataset_params.train_dataset_params, and
imagenet_renet50 recipe's dataset_params.train_dataloader_params).
The same logic holds for the validation set as well.
To demonstrate, let's look at what a configuration for training with the same data settings as in the previous code snippet looks like:
yamltrain_dataloader: imagenet_resnet50_train
val_dataloader: imagenet_resnet50_val
dataset_params:
train_dataset_params:
root: /my_data_dir/Imagenet/train
train_dataloader_params:
batch_size: 4
shuffle: True
val_dataset_params:
root: /my_data_dir/Imagenet/val
val_dataloader_params:
batch_size: 16
DataLoaders - Additional params
In addition to the parameters that are supported by the
torch.utils.data.DataLoader class, SuperGradients also provide additional parameters:*
min_samples - When present, this parameter will guarantee that at least min_samples items will be processed in each epoch. It is useful when working with small datasets.
To use this option, simply add this parameter to the dataloader_params dictionary, and set it to the desired value:
yamltrain_dataloader: imagenet_resnet50_train
dataset_params:
train_dataloader_params:
batch_size: 4
shuffle: True
min_samples: 1024
On the technical side, when this parameter is se, SuperGradients will attach the RandomSampler to the DataLoader, and set it's
num_samples parameter to min_samples.Using Custom Datasets
Suppose we already have our own
torch.utils.data.Dataset class:
pythonimport torchclass MyCustomDataset(torch.utils.data.Dataset):
def __init__(self, train: bool, image_size: int):
...
def __getitem__(self, item):
...
return inputs, targets # Or inputs, targets, additional_batch_items
#### A.
__getitem__
You need to make sure that the __getitem__ method of your dataset complies with the following format:
- inputs = batch_items[0] : model input - The type might depend on the model you are using.
- targets = batch_items[1] : Target that will be used to compute loss/metrics - The type might depend on the function you are using.
- [OPTIONAL] additional_batch_items = batch_items[2] : Dict made of any additional item that you might want to use.#### B. Train with your dataset
For coded training launch, we can instantiate it, then use it in the same way as the first code snippet to create
the data loaders and call train():
pythonfrom my_dataset import MyCustomDataset
from super_gradients.training import dataloaders, Trainertrain_dataset = MyCustomDataset(train=True, image_size=64)
valid_dataset = MyCustomDataset(train=False, image_size=128)
train_dataloader = dataloaders.get(dataset=train_dataset, dataloader_params={"batch_size": 4, "shuffle": True})
valid_dataloader = dataloaders.get(dataset=valid_dataset, dataloader_params={"batch_size": 16})
trainer = Trainer("my_custom_dataset_training_experiment")
model = ...
train_params = {...}
trainer.train(model=model, training_params=train_params, train_loader=train_dataloader, valid_loader=valid_dataloader)
Using Custom Datasets - Training with Configuration Files
When using configuration files, for example, training using train_from_recipe (or similar, when the underlying train method that is being called is Trainer.train_from_config(...)), In your
my_dataset.py, register your dataset class by decorating the class with register_dataset:
pythonimport torch
from super_gradients.common.registry.registry import register_dataset@register_dataset("my_custom_dataset")
class MyCustomDataset(torch.utils.data.Dataset):
def __init__(self, train: bool, image_size: int):
...
Then, use your newly registered dataset class in your configuration (of course, it can be split, use defaults, etc.) by referencing its name in the dataset entry inside dataloader_params while leaving out (or leaving empty) train_dataloader and valid_dataloader:
yamldataset_params:
train_dataset_params:
train: True
image_size: 64
train_dataloader_params:
dataset: my_custom_dataset
batch_size: 4
shuffle: True
val_dataset_params:
train: False
image_size: 128
val_dataloader_params:
dataset: my_custom_dataset
batch_size: 16
Last, in your
my_train_from_recipe_script.py file, import the newly registered class (even though the class itself is unused, just to trigger the registry):
pythonfrom omegaconf import DictConfig
import hydra
import pkg_resources
from my_dataset import MyCustomDataset
from super_gradients import Trainer, init_trainer
@hydra.main(config_path=pkg_resources.resource_filename("super_gradients.recipes", ""), version_base="1.2")
def main(cfg: DictConfig) -> None:
Trainer.train_from_config(cfg)
def run():
init_trainer()
main()
if __name__ == "__main__":
run()
Adding test datasets
In addition to the train and validation datasets, you can also add a test dataset or multiple test datasets to your configuration file.
At the end of training, metrics from each test dataset will be computed and returned in final results.
#### Single test dataset
To add a single test dataset to recipe, add following properties to your configuration file:
yamltest_dataloaders: <dataloader_name>dataset_params:
test_dataset_params:
...
test_dataloader_params:
...
#### Multiple test datasets
In addition to the train and validation datasets, you can also add a test dataset or multiple test datasets to your configuration file.
This is how you can achieve this using YAML file:
#### Explicitly specifying all parameters
yamltest_dataloaders:
test_dataset_name_1: <dataloader_name>
test_dataset_name_2: <dataloader_name>dataset_params:
test_dataset_params:
test_dataset_name_1:
...
test_dataset_name_2:
...
test_dataloader_params:
test_dataset_name_1:
...
test_dataset_name_2:
...
#### Without dataloader names
A
test_dataloaders property of the configuration file is optional and can be skipped.
You may want to use this option when you don't have a dataloader factory method registered.
In this case you have to specify a dataset class in corresponding dataloaders params.yamldataset_params:
test_dataset_params:
test_dataset_name_1:
...
test_dataset_name_2:
... test_dataloader_params:
test_dataset_name_1:
dataset: <dataset_class_name>
...
test_dataset_name_2:
dataset: <dataset_class_name>
...
#### Without dataloader params
A
dataset_params.test_dataloader_params property is optional and can be skipped.
In this case dataset_params.val_dataloader_params will be used for instantiating test dataloaders.
Please note that if you don't use test_dataloaders and test_dataloader_params properties, a dataset_params.val_dataloader_params
must contain a dataset property specifying class name of the dataset to use.yamldataset_params:
test_dataset_params:
test_dataset_name_1:
...
test_dataset_name_2:
...
---
Documentation/Source/Dataloader Adapter (documentation/source/dataloader_adapter.md)
Dataset Adapter
With diverse dataset structures available, ensuring compatibility with SuperGradients (SG) can be challenging. This is where the DataloaderAdapter plays a pivotal role. This tutorial takes you through the importance, implementation, and advantages of using the DataloaderAdapter with SG.
Why Dataset Adapter?
Datasets come in a myriad of structures. However, SG requires data in a specific format.
For instance, consider the Object Detection Format:
Image format should be: (BS, H, W, C) i.e., channel last.
Targets should be in the format: (BS, 6), where 6 represents (sample_id, class_id, label, cx, cy, w, h).
The overhead of adjusting each dataset manually can be cumbersome. Enter DataloaderAdapter β designed to automatically understand your dataset structure and mold it for SG compatibility.
Why Do We Need the Dataset Adapter?
While Datasets come in various structures and formats, SG expects data in a specific format to be able to run.
Example: Object Detection Format
- Image format: (BS, H, W, C) i.e. channel last
- Targets format: (BS, 6) where 6 represents (sample_id, class_id, label, cx, > cy, w, h).
This means that you should either use one of SuperGradient's built-in Dataset class if it supports your dataset structure, or, if your dataset is too custom for it, inherit from SG datasets and bring all the required changes.
While this is all right in most cases, it can be cumbersome when you just want to quickly experiment with a new dataset.
To reduce this overhead, SuperGradients introduced the concept of
DataloaderAdapter. Instead of requiring you to write all the transformations required to use SG, the DataloaderAdapter will infer anything possible directly from your data. Whenever something cannot be inferred with 100% confidence, you will be asked a question with all the required context for you to properly answer.Let's see this in practice with an example. Let's start with
SBDataset datasetExemple 1 - Segmentation Adapter on
SBDataset DatasetIn this section, we'll walk through the process of preparing the
SBDataset dataset for use in SuperGradients. We'll highlight the challenges and demonstrate how the Adapter can simplify the process.
1. Preparing the Dataset without Adapter
pythonfrom torchvision.datasets import SBDatasettry:
# There is a bug with
torchvision.datasets.SBDataset that raises RuntimeError after downloading, so we just ignore it
SBDataset(root="data", mode='segmentation', download=True)
except RuntimeError:
pass
Downloading https://www2.eecs.berkeley.edu/Research/Projects/CS/vision/grouping/semantic_contours/benchmark.tgz to data/benchmark.tgz
100%|ββββββββββ| 1419539633/1419539633 [00:32<00:00, 43301819.66it/s]
Extracting data/benchmark.tgz to data
Downloading https://www.cs.cornell.edu/~bharathh/ to data/train_noval.txt
20563it [00:00, 1012436.88it/s]
pythonfrom torchvision.transforms import Compose, ToTensor, Resize, InterpolationMode
transforms = Compose([ToTensor(), Resize((512, 512), InterpolationMode.NEAREST)])
def sample_transform(image, mask):
return transforms(image), transforms(mask)
train_set = SBDataset(root="data", mode='segmentation', download=False, transforms=sample_transform)
Now let's see what we get when instantiating a
Dataloader
pythonfrom torch.utils.data import DataLoadertrain_loader = DataLoader(train_set, batch_size=20, shuffle=True)
_images, labels = next(iter(train_loader))
labels.unique()
tensor([0.0000, 0.0118, 0.0157, 0.0196, 0.0235, 0.0275, 0.0353, 0.0431, 0.0471,
0.0549, 0.0588, 0.0627, 0.0706, 0.0745, 0.0784])
As you can see, the labels are normalized (0-1). This is all right, but it is not the format expected by SuperGradients.
Let's now see how the Adapter helps.
2. Introducing Adapter
The Adapter helps us skip manual data preparations and dives right into creating a dataloader that SuperGradients expects.
pythonfrom super_gradients.training.dataloaders.adapters import SegmentationDataloaderAdapterFactorytrain_loader = SegmentationDataloaderAdapterFactory.from_dataset(dataset=train_set, batch_size=20, shuffle=True, config_path='local_cache.json')
_images, labels = next(iter(train_loader))
labels.unique()
[2023-10-29 15:25:36] INFO - data_config.py - Cache deactivated for
SegmentationDataConfig.
--------------------------------------------------------------------------------
How many classes does your dataset include?
--------------------------------------------------------------------------------
Enter your response >>> 21
Great! You chose:
21
--------------------------------------------------------------------------------
Does your dataset provide a batch or a single sample?
--------------------------------------------------------------------------------
- Image shape: torch.Size([3, 512, 512])
- Mask shape: torch.Size([1, 512, 512])
Options:
[0] | Batch of Samples (e.g. torch Dataloader)
[1] | Single Sample (e.g. torch Dataset)
Your selection (Enter the corresponding number) >>> 1
Great! You chose: Single Sample (e.g. torch Dataset)
--------------------------------------------------------------------------------
In which format are your images loaded ?
--------------------------------------------------------------------------------
Options:
[0] | RGB
[1] | BGR
[2] | LAB
[3] | Other
Your selection (Enter the corresponding number) >>> 0
Great! You chose: RGB tensor([ 0, 1, 2, 3, 4, 7, 8, 9, 12, 13, 15, 16, 18, 19, 20])
You can see that the mask is now encoded as
int, which is the representation used in SuperGradients.It's important to note that the dataset adapter also support different dataset format such as one hot, ensuring that the output (
labels here) is in the right format to use within SuperGradients.Example II - Detection Adapter on a Dictionary based Dataset
Some datasets return a more complex data structure than the previous example.
For instance, the
COCO dataset implementation from pytorch returns a list of dictionaries representing the labels.Let's have a look:
pythonDownload the zip file
!wget https://deci-pretrained-models.s3.amazonaws.com/coco2017_small.zipUnzip the downloaded file
!unzip coco2017_small.zip > /dev/null
--2023-10-29 15:27:31-- https://deci-pretrained-models.s3.amazonaws.com/coco2017_small.zip
Resolving deci-pretrained-models.s3.amazonaws.com (deci-pretrained-models.s3.amazonaws.com)... 54.231.134.129, 52.217.71.68, 52.217.138.65, ...
Connecting to deci-pretrained-models.s3.amazonaws.com (deci-pretrained-models.s3.amazonaws.com)|54.231.134.129|:443... connected.
HTTP request sent, awaiting response... 200 OK
Length: 246116231 (235M) [application/zip]
Saving to: βcoco2017_small.zipβ
coco2017_small.zip 100%[===================>] 234.71M 38.7MB/s in 6.7s
2023-10-29 15:27:38 (34.9 MB/s) - βcoco2017_small.zipβ saved [246116231/246116231]
pythonfrom torchvision.datasets import CocoDetection
from torchvision.transforms import Compose, ToTensor, Resize, InterpolationMode
from torchvision.datasets import SBDataset
image_transform = Compose([ToTensor(), Resize((512, 512))])
train_set = CocoDetection(root='coco2017_small/images/train2017', annFile='coco2017_small/annotations/instances_train2017.json', transform=image_transform)
train_set = CocoDetection(root='coco2017_small/images/val2017', annFile='coco2017_small/annotations/instances_val2017.json', transform=image_transform)
image, targets = next(iter(train_set))
loading annotations into memory...
Done (t=0.10s)
creating index...
index created!
loading annotations into memory...
Done (t=0.05s)
creating index...
index created!
pythonprint(f"Number of targets: {len(targets)}, First target structure: {targets[0]}")
Observe the dataset output's nested dictionary structure? This complicates things for the Dataset Adapter as it's unsure which fields detail the bounding box.
To solve this, we utilize an extractor function.
#### The Extractor's Role
Simply put, the extractor translates your dataset's output into a format the Adapter understands. For our dataset, it will take the image and annotations, then return the bounding box data, including the label and coordinates.
Worried about bounding box format like
xyxy_label or label_xywh? Don't be. The Adapter is designed to recognize them.For further guidance on extractor functions, see the official documentation.
pythonimport torchdef coco_labels_extractor(sample) -> torch.Tensor:
_, annotations = sample # annotations = [{"bbox": [1.08, 187.69, 611.59, 285.84], "category_id": 51}, ...]
labels = []
for annotation in annotations:
class_id = annotation["category_id"]
bbox = annotation["bbox"]
labels.append((class_id, *bbox))
return torch.Tensor(labels) # torch.Tensor([[51, 1.08, 187.69, 611.59, 285.84], ...])
coco_labels_extractor(sample=next(iter(train_set)))
tensor([[ 64.0000, 236.9800, 142.5100, 24.7000, 69.5000],
[ 72.0000, 7.0300, 167.7600, 149.3200, 94.8700],
[ 72.0000, 557.2100, 209.1900, 81.3500, 78.7300],
[ 62.0000, 358.9800, 218.0500, 56.0000, 102.8300],
[ 62.0000, 290.6900, 218.0000, 61.8300, 98.4800],
[ 62.0000, 413.2000, 223.0100, 30.1700, 81.3600],
[ 62.0000, 317.4000, 219.2400, 21.5800, 11.5900],
[ 1.0000, 412.8000, 157.6100, 53.0500, 138.0100],
[ 1.0000, 384.4300, 172.2100, 15.1200, 35.7400],
[ 78.0000, 512.2200, 205.7500, 14.7400, 15.9700],
[ 82.0000, 493.1000, 174.3400, 20.2900, 108.3100],
[ 84.0000, 604.7700, 305.8900, 14.3400, 45.7100],
[ 84.0000, 613.2400, 308.2400, 12.8800, 46.4400],
[ 85.0000, 447.7700, 121.1200, 13.9700, 21.8800],
[ 86.0000, 549.0600, 309.4300, 36.6800, 89.6700],
[ 86.0000, 350.7600, 208.8400, 11.3700, 22.5500],
[ 62.0000, 412.2500, 219.0200, 9.6300, 12.5200],
[ 86.0000, 241.2400, 194.9900, 14.2200, 17.6300],
[ 86.0000, 336.7900, 199.5000, 9.7300, 16.7300],
[ 67.0000, 321.2100, 231.2200, 125.5600, 88.9300]])
This output is all you need to get started. Now we can use the Dataloader Adapters!
pythonfrom super_gradients.training.dataloaders.adapters import DetectionDataloaderAdapterFactory
from data_gradients.dataset_adapters.config.data_config import DetectionDataConfig
adapter_config = DetectionDataConfig(labels_extractor=coco_labels_extractor, cache_path="coco_adapter_cache.json")
train_loader = DetectionDataloaderAdapterFactory.from_dataset(
dataset=train_set,
config=adapter_config,
batch_size=5,
drop_last=True,
)
val_loader = DetectionDataloaderAdapterFactory.from_dataset(
dataset=train_set,
config=adapter_config,
batch_size=5,
drop_last=True,
)
/usr/local/lib/python3.10/dist-packages/torchvision/transforms/functional.py:1603: UserWarning: The default value of the antialias parameter of all the resizing transforms (Resize(), RandomResizedCrop(), etc.) will change from None to True in v0.17, in order to be consistent across the PIL and Tensor backends. To suppress this warning, directly pass antialias=True (recommended, future default), antialias=None (current default, which means False for Tensors and True for PIL), or antialias=False (only works on Tensors - PIL will still use antialiasing). This also applies if you are using the inference transforms from the models weights: update the call to weights.transforms(antialias=True).
warnings.warn(
[2023-10-29 15:27:41] INFO - data_config.py - Cache deactivated for
DetectionDataConfig.
[2023-10-29 15:27:41] INFO - detection_adapter_collate_fn.py - You are using Detection Adapter. Please note that it was designed specifically for YOLONAS, YOLOX and PPYOLOE.
Number of targets: 20, First target structure: {'segmentation': [[240.86, 211.31, 240.16, 197.19, 236.98, 192.26, 237.34, 187.67, 245.8, 188.02, 243.33, 176.02, 250.39, 186.96, 251.8, 166.85, 255.33, 142.51, 253.21, 190.49, 261.68, 183.08, 258.86, 191.2, 260.98, 206.37, 254.63, 199.66, 252.51, 201.78, 251.8, 212.01]], 'area': 531.8071000000001, 'iscrowd': 0, 'image_id': 139, 'bbox': [236.98, 142.51, 24.7, 69.5], 'category_id': 64, 'id': 26547}
--------------------------------------------------------------------------------
How many classes does your dataset include?
--------------------------------------------------------------------------------
Enter your response >>> 80
Great! You chose:
80
--------------------------------------------------------------------------------
In which format are your images loaded ?
--------------------------------------------------------------------------------
Options:
[0] | RGB
[1] | BGR
[2] | LAB
[3] | Other
Your selection (Enter the corresponding number) >>> 0
Great! You chose: RGB
--------------------------------------------------------------------------------
Which comes first in your annotations, the class id or the bounding box?
--------------------------------------------------------------------------------
Here's a sample of how your labels look like:
Each line corresponds to a bounding box.
tensor([[ 64.0000, 236.9800, 142.5100, 24.7000, 69.5000],
[ 72.0000, 7.0300, 167.7600, 149.3200, 94.8700],
[ 72.0000, 557.2100, 209.1900, 81.3500, 78.7300],
[ 62.0000, 358.9800, 218.0500, 56.0000, 102.8300]])
Options:
[0] | Label comes first (e.g. [class_id, x1, y1, x2, y2])
[1] | Bounding box comes first (e.g. [x1, y1, x2, y2, class_id])
Your selection (Enter the corresponding number) >>> 0
Great! You chose: Label comes first (e.g. [class_id, x1, y1, x2, y2])
--------------------------------------------------------------------------------
What is the bounding box format?
--------------------------------------------------------------------------------
Here's a sample of how your labels look like:
Each line corresponds to a bounding box.
tensor([[ 64.0000, 236.9800, 142.5100, 24.7000, 69.5000],
[ 72.0000, 7.0300, 167.7600, 149.3200, 94.8700],
[ 72.0000, 557.2100, 209.1900, 81.3500, 78.7300],
[ 62.0000, 358.9800, 218.0500, 56.0000, 102.8300]])
Options:
[0] | xyxy: x-left, y-top, x-right, y-bottom (Pascal-VOC format)
[1] | xywh: x-left, y-top, width, height (COCO format)
[2] | cxcywh: x-center, y-center, width, height (YOLO format)
Your selection (Enter the corresponding number) >>> 1
[2023-10-29 15:28:40] INFO - detection_adapter_collate_fn.py - You are using Detection Adapter. Please note that it was designed specifically for YOLONAS, YOLOX and PPYOLOE.
python_image, targets = next(iter(train_loader))
pythontargets.shape # [N, 6] format with 6 representing (sample_id, class_id, cx, cy, w, h)
torch.Size([22, 6])
pythontargets[:3]
tensor([[ 0.0000, 64.0000, 249.3300, 177.2600, 24.7000, 69.5000],
[ 0.0000, 72.0000, 81.6900, 215.1950, 149.3200, 94.8700],
[ 0.0000, 72.0000, 597.8850, 248.5550, 81.3500, 78.7300]])
III. Use your Adapted Dataloader to train a model
Now that we have an adapter for a detection dataset, let's use it to launch a training of YoloNAS on it!
This is of course for the sake of the example, since YoloNAS was originally trained using the SuperGradients implementation of COCO Dataset. You can replace the
COCO dataset with any of your dataset.pythonfrom omegaconf import OmegaConf
from hydra.utils import instantiatefrom super_gradients import Trainer
from super_gradients.training import models
from super_gradients.common.object_names import Models
from super_gradients.training import training_hyperparams
from super_gradients.common.environment.cfg_utils import load_recipe
trainer = Trainer(experiment_name="yolonas_training_with_adapter", ckpt_root_dir="../../scripts/")
model = models.get(model_name=Models.YOLO_NAS_S, num_classes=adapter_config.n_classes, pretrained_weights="coco")
yolonas_recipe = load_recipe(config_name="coco2017_yolo_nas_s",
overrides=[f"arch_params.num_classes={adapter_config.n_classes}",
"training_hyperparams.max_epochs=1",
"training_hyperparams.mixed_precision=False"])
yolonas_recipe = OmegaConf.to_container(instantiate(yolonas_recipe))
training_params = yolonas_recipe['training_hyperparams']
trainer.train(model=model, training_params=training_params, train_loader=train_loader, valid_loader=val_loader)
[2023-10-29 15:29:42] INFO - checkpoint_utils.py - License Notification: YOLO-NAS pre-trained weights are subjected to the specific license terms and conditions detailed in
https://github.com/Deci-AI/super-gradients/blob/master/LICENSE.YOLONAS.md
By downloading the pre-trained weight files you agree to comply with these terms.
Downloading: "https://sghub.deci.ai/models/yolo_nas_s_coco.pth" to /root/.cache/torch/hub/checkpoints/yolo_nas_s_coco.pth
100%|ββββββββββ| 73.1M/73.1M [00:00<00:00, 81.0MB/s]
[2023-10-29 15:29:43] INFO - checkpoint_utils.py - Successfully loaded pretrained weights for architecture yolo_nas_s
[2023-10-29 15:29:44] INFO - sg_trainer.py - Starting a new run with
run_id=RUN_20231029_152944_310569
[2023-10-29 15:29:44] INFO - sg_trainer.py - Checkpoints directory: ./yolonas_training_with_adapter/RUN_20231029_152944_310569
[2023-10-29 15:29:44] INFO - sg_trainer.py - Using EMA with params {'decay': 0.9997, 'decay_type': 'threshold', 'beta': 15}
The console stream is now moved to ./yolonas_training_with_adapter/RUN_20231029_152944_310569/console_Oct29_15_29_44.txt
[2023-10-29 15:29:45] WARNING - callbacks.py - Number of warmup steps (1000) is greater than number of steps in epoch (100). Warmup steps will be capped to number of steps in epoch to avoid interfering with any pre-epoch LR schedulers.
/usr/local/lib/python3.10/dist-packages/torchvision/transforms/functional.py:1603: UserWarning: The default value of the antialias parameter of all the resizing transforms (Resize(), RandomResizedCrop(), etc.) will change from None to True in v0.17, in order to be consistent across the PIL and Tensor backends. To suppress this warning, directly pass antialias=True (recommended, future default), antialias=None (current default, which means False for Tensors and True for PIL), or antialias=False (only works on Tensors - PIL will still use antialiasing). This also applies if you are using the inference transforms from the models weights: update the call to weights.transforms(antialias=True).
warnings.warn(
/usr/local/lib/python3.10/dist-packages/super_gradients/training/utils/collate_fn/detection_collate_fn.py:29: UserWarning: To copy construct from a tensor, it is recommended to use sourceTensor.clone().detach() or sourceTensor.clone().detach().requires_grad_(True), rather than torch.tensor(sourceTensor).
images_batch = [torch.tensor(img) for img in images_batch]
/usr/local/lib/python3.10/dist-packages/super_gradients/training/utils/collate_fn/detection_collate_fn.py:43: UserWarning: To copy construct from a tensor, it is recommended to use sourceTensor.clone().detach() or sourceTensor.clone().detach().requires_grad_(True), rather than torch.tensor(sourceTensor).
labels_batch = [torch.tensor(labels) for labels in labels_batch]
[2023-10-29 15:29:45] INFO - sg_trainer_utils.py - TRAINING PARAMETERS:
- Mode: Single GPU
- Number of GPUs: 0 (0 available on the machine)
- Full dataset size: 500 (len(train_set))
- Batch size per GPU: 5 (batch_size)
- Batch Accumulate: 1 (batch_accumulate)
- Total batch size: 5 (num_gpus * batch_size)
- Effective Batch size: 5 (num_gpus batch_size batch_accumulate)
- Iterations per epoch: 100 (len(train_loader))
- Gradient updates per epoch: 100 (len(train_loader) / batch_accumulate)
[2023-10-29 15:29:45] INFO - sg_trainer.py - Started training for 1 epochs (0/0)
Train epoch 0: 100%|ββββββββββ| 100/100 [20:18<00:00, 12.18s/it, PPYoloELoss/loss=4, PPYoloELoss/loss_cls=1.72, PPYoloELoss/loss_dfl=2.2, PPYoloELoss/loss_iou=0.472, gpu_mem=0]
Validating: 100%|ββββββββββ| 100/100 [06:34<00:00, 3.94s/it]
[2023-10-29 15:56:43] INFO - base_sg_logger.py - Checkpoint saved in ./yolonas_training_with_adapter/RUN_20231029_152944_310569/ckpt_best.pth
[2023-10-29 15:56:43] INFO - sg_trainer.py - Best checkpoint overriden: validation [email protected]:0.95: 0.0005365016404539347
===========================================================
SUMMARY OF EPOCH 0
βββ Train
β βββ Ppyoloeloss/loss_cls = 1.7168
β βββ Ppyoloeloss/loss_iou = 0.4717
β βββ Ppyoloeloss/loss_dfl = 2.2035
β βββ Ppyoloeloss/loss = 3.9977
βββ Validation
βββ Ppyoloeloss/loss_cls = 2.4528
βββ Ppyoloeloss/loss_iou = 0.5016
βββ Ppyoloeloss/loss_dfl = 2.2003
βββ Ppyoloeloss/loss = 4.807
βββ [email protected]:0.95 = 0.0052
βββ [email protected]:0.95 = 0.007
βββ [email protected]:0.95 = 0.0005
βββ [email protected]:0.95 = 0.0007
===========================================================
[2023-10-29 15:56:45] INFO - sg_trainer.py - RUNNING ADDITIONAL TEST ON THE AVERAGED MODEL...
Validating epoch 1: 100%|ββββββββββ| 100/100 [06:33<00:00, 3.93s/it]
IV. Dig deeper into the Adapter
By default, any parameter that could not be confidently infered will trigger a question.
But you have the possibility to set these parameters in advance through the config object. In the previous example we had to set
labels_extractor explicitly. Now let's set all the parameters
pythonfrom super_gradients.training.dataloaders.adapters import DetectionDataloaderAdapterFactory
from data_gradients.dataset_adapters.config.data_config import DetectionDataConfig
from data_gradients.utils.data_classes.image_channels import ImageChannels
class_names = [category['name'] for category in train_set.coco.loadCats(train_set.coco.getCatIds())]adapter_config = DetectionDataConfig(
labels_extractor=coco_labels_extractor,
is_label_first=True,
class_names=class_names,
image_channels=ImageChannels.from_str("RGB"),
xyxy_converter='xywh',
cache_path="coco_adapter_cache_with_default.json"
)
This can now be used and you don't need to answer any question
pythontrain_loader = DetectionDataloaderAdapterFactory.from_dataset(
dataset=train_set,
config=adapter_config,
batch_size=5,
drop_last=True,
)
val_loader = DetectionDataloaderAdapterFactory.from_dataset(
dataset=train_set,
config=adapter_config,
batch_size=5,
drop_last=True,
)_image, targets = next(iter(train_loader))
print(targets.shape) # [N, 6] format with 6 representing (sample_id, class_id, cx, cy, w, h)
[2023-10-29 16:15:09] INFO - detection_adapter_collate_fn.py - You are using Detection Adapter. Please note that it was designed specifically for YOLONAS, YOLOX and PPYOLOE.
/usr/local/lib/python3.10/dist-packages/torchvision/transforms/functional.py:1603: UserWarning: The default value of the antialias parameter of all the resizing transforms (Resize(), RandomResizedCrop(), etc.) will change from None to True in v0.17, in order to be consistent across the PIL and Tensor backends. To suppress this warning, directly pass antialias=True (recommended, future default), antialias=None (current default, which means False for Tensors and True for PIL), or antialias=False (only works on Tensors - PIL will still use antialiasing). This also applies if you are using the inference transforms from the models weights: update the call to weights.transforms(antialias=True).
warnings.warn(
[2023-10-29 16:15:09] INFO - detection_adapter_collate_fn.py - You are using Detection Adapter. Please note that it was designed specifically for YOLONAS, YOLOX and PPYOLOE.
/usr/local/lib/python3.10/dist-packages/super_gradients/training/utils/collate_fn/detection_collate_fn.py:29: UserWarning: To copy construct from a tensor, it is recommended to use sourceTensor.clone().detach() or sourceTensor.clone().detach().requires_grad_(True), rather than torch.tensor(sourceTensor).
images_batch = [torch.tensor(img) for img in images_batch]
/usr/local/lib/python3.10/dist-packages/super_gradients/training/utils/collate_fn/detection_collate_fn.py:43: UserWarning: To copy construct from a tensor, it is recommended to use sourceTensor.clone().detach() or sourceTensor.clone().detach().requires_grad_(True), rather than torch.tensor(sourceTensor).
labels_batch = [torch.tensor(labels) for labels in labels_batch]
Load from existing cache
You can use the cache of an adapter you already used in the past. This will allow you skip the questions that were already asked in the previous run.
pythonThe new config will load the answer from questions asked in the previous run.
adapter_config = DetectionDataConfig(
labels_extractor=coco_labels_extractor,
cache_path="coco_adapter_cache_with_default.json" # Name of the previous cache
)train_loader = DetectionDataloaderAdapterFactory.from_dataset(
dataset=train_set,
config=adapter_config,
batch_size=5,
drop_last=True,
)
val_loader = DetectionDataloaderAdapterFactory.from_dataset(
dataset=train_set,
config=adapter_config,
batch_size=5,
drop_last=True,
)
_image, targets = next(iter(train_loader))
pythontargets.shape # [N, 6] format with 6 representing (sample_id, class_id, cx, cy, w, h)
torch.Size([22, 6])
As you can see, no question was asked and we still get the targets adapted into the SuperGradients format.
---
Documentation/Source/DetectionOutputAdapter (documentation/source/DetectionOutputAdapter.md)
DetectionOutputAdapter
The DetectionOutputAdapter is a class that converts the output of a detection model into a user-appropriate format.
For instance, it can be used to convert the format of bounding boxes from CYXHW to XYXY, or to change the layout of the elements
in the output tensor from [X1, Y1, X2, Y2, Confidence, Class] to [Class, Confidence, X1, Y1, X2, Y2].
Features
* Easy rearrangement of the elements in the output tensor
* Easy conversion of the bounding box format
* Support of JIT Tracing & Scripting
* Support of ONNX export
Usage
We start by introducing the concept of a
format. A format represents a specific layout of the elements in the output tensor.
Currently, there is only one type of formats supported - ConcatenatedTensorFormat which represents a layout where all predictions concatenated into a single tensor.
Additional formats can be added in the future (Like DictionaryOfTensorsFormat).ConcatenatedTensorFormat requires that input is a tensor and has the following shape:
* Tensor of shape [N, Elements] - N is the number of predictions, Elements is the concatenated vector of attributes per box.
* Tensor of shape [B, N, Elements] - B is the batch dimension, N and Elements as above. To instantiate the
DetectionOutputAdapter we have to describe the input and output formats for our predictions:Let's imagine model emits predictions in the following format:
python[N, 10] (cx, cy, w, h, class, confidence, attributes..)
example_input = [
# cx cy w h class, confidence, attribute a, attribute b, attribute c, attribute d
[0.465625, 0.5625, 0.13125, 0.125, 0, 0.968, 0.350, 0.643, 0.640, 0.453],
[0.103125, 0.1671875, 0.10625, 0.134375, 1, 0.897, 0.765, 0.654, 0.324, 0.816],
[0.078125, 0.078125, 0.15625, 0.15625, 2., 0.423, 0.792, 0.203, 0.653, 0.777],
...
]
The corresponding format definition would look like this:
pythonfrom super_gradients.training.datasets.data_formats import ConcatenatedTensorFormat, BoundingBoxesTensorSliceItem, TensorSliceItem, NormalizedCXCYWHCoordinateFormatinput_format = ConcatenatedTensorFormat(
layout=(
BoundingBoxesTensorSliceItem(name="bboxes", format=NormalizedCXCYWHCoordinateFormat()),
TensorSliceItem(name="class", length=1),
TensorSliceItem(name="confidence", length=1),
TensorSliceItem(name="attributes", length=4),
)
)
For sake of demonstration, let's assume that we want to convert the output to the following format:
python[N, 10] (class, attributes, x1, y1, x2, y2)
[
# class, attribute a, attribute b, attribute c, attribute d, x1, y1, x2, y2
[ 0, 0.350, 0.643, 0.640, 0.453, 256, 320, 340, 400],
[ 1, 0.765, 0.654, 0.324, 0.816, 32, 64, 100, 150],
[ 2, 0.792, 0.203, 0.653, 0.777, 0, 0, 100, 100],
...
]
* The
class and attributes are the same as in the input format but comes first
* The format of bounding boxes is changed from NormalizedCXCYWHCoordinateFormat to XYXYCoordinateFormat
* The confidence is removed from the outputThe corresponding format definition would look like this:
pythonfrom super_gradients.training.datasets.data_formats import ConcatenatedTensorFormat, BoundingBoxesTensorSliceItem, TensorSliceItem, XYXYCoordinateFormatoutput_format = ConcatenatedTensorFormat(
layout=(
TensorSliceItem(name="class", length=1),
TensorSliceItem(name="attributes", length=4),
BoundingBoxesTensorSliceItem(name="bboxes", format=XYXYCoordinateFormat()),
)
)
Now we can construct the
DetectionOutputAdapter and attach it to the model:pythonfrom super_gradients.training.datasets.data_formats import DetectionOutputAdapteroutput_adapter = DetectionOutputAdapter(input_format, output_format, image_shape=(640,640))
model = nn.Sequential(
create_model(),
create_nms(),
output_adapter
)
To test how the output adapter transforms dummy input one can easily run it alone:
pythonoutput = output_adapter(torch.from_numpy(example_input)).numpy()
print(output)Prints:
[
# class, attribute a, attribute b, attribute c, attribute d, x1, y1, x2, y2
[ 0, 0.350, 0.643, 0.640, 0.453, 256, 320, 340, 400],
[ 1, 0.765, 0.654, 0.324, 0.816, 32, 64, 100, 150],
[ 2, 0.792, 0.203, 0.653, 0.777, 0, 0, 100, 100]
]
Not supported features
Currently
DetectionOutputAdapter does not support the following features:*
argmax operation over a slice of confidences for [C] classes (Useful to compute argmax(class confidences))
Multiplication of two slices (Useful to compute confidence class)---
Documentation/Source/Device (documentation/source/device.md)
Training Modes
SuperGradients allows users to train models on different modes:
1. CPU
2. single GPU - (CUDA)
3. multiple GPUs' - Data Parallel (DP)
4. multiple GPUs' - Distributed Data Parallel (DDP)
1. CPU
Requirement: None.How to use it: If you don't have any CUDA device available, your training will automatically be run on CPU.
Otherwise, the default device will be CUDA, but you can still easily set it to CPU using
setup_device as follow:
pyfrom super_gradients import Trainer
from super_gradients.training.utils.distributed_training_utils import setup_devicesetup_device(device='cpu')
Unchanged
trainer = Trainer(...)
trainer.train(...)
2. CUDA
Requirement: Having at least one CUDA device availableHow to use it: If you have at least one CUDA device, nothing! Otherwise, you will have to use CPU...
3. DP - Data Parallel
DataParallel (DP) is a single-process, multi-thread technic for scaling deep learning model training across multiple GPUs on a single machine.The general flow is as below
- Split the data into smaller chunks (mini-batch) on GPU:0
- Move one chunk of data per GPU
- Copy the model to all available GPUs
- Perform the forward pass on each GPU in parallel
- Gather the outputs on GPU:0
- Compute the loss on GPU:0
- Share the loss to all the GPUs
- Compute the gradients on each GPU
- Gather and sum up gradients on GPU:0
- Update model on GPU:0
For more detailed information, feel free to check out this blog for a more in-depth explanation.
Requirement: Having at least one CUDA devices available
How to use it: All you need to do is to call a magic function
setup_device before instantiating the Trainer.
pyfrom super_gradients import Trainer
from super_gradients.training.utils.distributed_training_utils import setup_deviceLaunch DP on 4 GPUs'
setup_device(multi_gpu='DP', num_gpus=4)Unchanged
trainer = Trainer(...)
trainer.train(...)
Tip: To optimize runtime we recommend to call setup_device as early as possible.4. DDP - Distributed Data Parallel
Distributed Data Parallel (DDP) is a powerful technique for scaling deep learning model training across multiple GPUs.
It involves the use of multiple processes, each running on a different GPU and having its own instance of the model.
The processes communicate only to exchange gradients, making it a highly efficient and more scalable solution for training large models than Data Parallel (DP).Although DDP can be more complex to set up than DP, the SuperGradients library abstracts away the complexity by handling the setup process behind the scenes.
This makes it easy for users to take advantage of the benefits of DDP without having to worry about technical details.
We highly recommend using DDP over DP whenever possible.
For more detailed information, feel free to check out this blog for a more in-depth explanation.
Requirement: Having multiple CUDA devices available
How to use it: All you need to do is to call a magic function
setup_device before instantiating the Trainer.
pyfrom super_gradients import Trainer
from super_gradients.training.utils.distributed_training_utils import setup_deviceLaunch DDP on 4 GPUs'
setup_device(num_gpus=4) # Equivalent to: setup_device(multi_gpu='DDP', num_gpus=4)Unchanged
trainer = Trainer(...)
trainer.train(...)
Tip: To optimize runtime we recommend to call setup_device as early as possible.
What should you be aware of when using DDP ?
#### A. DDP runs multiple processes
When running DDP, you will work with multiple processes that will go through the whole training loop.
This means that if you run DDP on 4 gpus, any action that you do will be run 4 times.
This impacts especially printing, logging, and file writing. To face this issue, SuperGradients provides a decorator
that will ensure that only one process will execute a specific function, whether you work with CPU, GPU, DP, or DDP.
In the following example, the
print_hello function will print Hello world only once, when it would be printed 4 times without the decorator...
pyfrom super_gradients.training.utils.distributed_training_utils import setup_device
from super_gradients.common.environment.ddp_utils import multi_process_safesetup_device(num_gpus=4)
@multi_process_safe # Try with and without this decorator
def print_hello():
print('Hello world')
print_hello()
#### B. DDP requires specific Metric implementation!
As explained, multiple processes are used to train a model with DDP, each on its own GPU.
This means that the metrics must be computed and aggregated across all the processes, and it requires the metric to be implemented using states.
States are attributes to be reduced. They are defined using the built-in method
add_state() and enable broadcasting
of the states among the different ranks when calling the compute() method.An example of a state would be the number of correct predictions, which will be summed across the different processes, and broadcasted to all of
them before computing the metric value. You can see an example below.
Feel free to check torchmetrics documentation for more information on how to implement your own metric.
Example
In the following example, we start with a custom metric implemented to run on a single device:
pyimport torch
from torchmetrics import Metric
class Top5Accuracy(Metric):
def __init__(self):
super().__init__()
self.correct = torch.tensor(0.)
self.total = torch.tensor(0.)
def update(self, preds: torch.Tensor, target: torch.Tensor):
batch_size = target.size(0)
# Get the top k predictions
_, pred = preds.topk(5, 1, True, True)
pred = pred.t()
# Count the number of correct predictions only for the highest 5
correct = pred.eq(target.view(1, -1).expand_as(pred))
correct5 = correct[:5].reshape(-1).float().sum(0)
self.correct += correct5.cpu()
self.total += batch_size
def compute(self):
return self.correct.float() / self.total
All you need to change to use your metric on DDP is to define your attributes self.correct and self.total with add_state and to define
a reduce function dist_reduce_fx that will be used to know how to combine the states when calling compute:
pyimport torch
import torchmetricsclass DDPTop1Accuracy(torchmetrics.Metric):
def __init__(self, dist_sync_on_step=False):
super().__init__(dist_sync_on_step=dist_sync_on_step)
self.add_state("correct", default=torch.tensor(0.), dist_reduce_fx="sum") # Set correct to be a state
self.add_state("total", default=torch.tensor(0), dist_reduce_fx="sum") # Set total to be a state
def update(self, preds: torch.Tensor, target: torch.Tensor):
batch_size = target.size(0)
# Get the top k predictions
_, pred = preds.topk(5, 1, True, True)
pred = pred.t()
# Count the number of correct predictions only for the highest 5
correct = pred.eq(target.view(1, -1).expand_as(pred))
correct5 = correct[:5].reshape(-1).float().sum(0)
self.correct += correct5
self.total += batch_size
def compute(self):
return self.correct.float() / self.total
Step by step explanation
1. DDP launches, and pytorch creates a separate instance of your custom metric for each process.
2. The update() method modifies the internal state of each instance in each process based on the inputs preds and target specific to that process.
3. After an epoch, each process will have a unique state, for example:
- Process 1: correct=50, total=100
- Process 2: correct=30, total=100
- Process 3: correct=100, total=100
4. Calling compute() triggers torchmetrics.Metric to gather and combine the states of each process. This reduction step can be customized by setting the dist_reduce_fx, which in this case is the sum. This usually happens at the end of the epoch.
- All processes: correct=180, total=300
5. The compute() method then calculates the metric value according to your implementation. In this example, every process will return the same result: 0.6 (180 correct predictions out of 300 total predictions).
6. Finally, calling reset() will reset the internal state of the metric, making it ready to accumulate new data at the start of the next epoch.C. When using DDP you may want to scale the learning rate
Using N GPUs in DDP mode, has an effect of increasing batch size by a factor of N.
And it has been shown that it may be necessary to scale the learning rate accordingly.
The rule of thumb is that if batch size is increased by a factor of N (Or N nodes used in DDP), the learning rate should be
also increased by a factor of N.
However, when it comes to adaptive optimizers like Adam, the situation is a bit different.
Adaptive optimizers like Adam automatically adjust the learning rate for each parameter based on the historical
gradient information.
They inherently adapt to the scale of the gradients and don't require manual adjustments of the learning rate
in the same way as fixed learning rate methods like SGD.
That being said, we still recommend to try out different learning rates to see the impact on the final metrics.
You can run experiments manually or use Hydra sweep syntax to run experiments with custom learning rates as follows:
bashpython -m super_gradients.train_from_recipe -m --config-name=coco2017_yolo_nas_s training_hyperparams.initial_lr=1e-3,5e-3,1e-4
---
How to set training mode with recipes ?
When using recipes you simply need to set values of gpu_mode and num_gpus.yamltraining_recipe.yaml
default:
- ......
Simply add this to run DDP on 4 nodes.
gpu_mode: DDP
num_gpus: 4
---
Documentation/Source/EMA (documentation/source/EMA.md)
Exponential Moving Average (EMA)
Exponential Moving Average or EMA is a technique used during training to smooth the noise in the training process and improve the generalization of the model.
It is a simple technique that can be used with any model and optimizer. Here's a recap how EMA works:
- At the start of training, the model parameters are copied to the EMA parameters.
- At each gradient update step, the EMA parameters are updated using the following formula:
pythonema_param = ema_param decay + param (1 - decay)
- On start of validation epoch the model parameters are replaced by the EMA parameters and reverted back on the end of validation epoch.
- At the end of training, the model parameters are replaced by the EMA parameters.
To enable use of EMA is SuperGradients one should add following parameters to the
training_params:pyfrom super_gradients import Trainer
trainer = Trainer(...)
trainer.train(
training_params={"ema": True, "ema_params": {"decay": 0.9999, "decay_type": "constant"}, ...},
...
)
The
decay is a hyperparameter that controls the speed of the EMA update. It's value must be in (0,1) range.
Larger values of decay will result in slower EMA model update.It is usually beneficial to have smaller decay values at the start of training and increase it as the training progresses.
In SuperGradients we support several types of changing decay value over time:
-
constant: "ema_params": {"decay": 0.9999, "decay_type": "constant"}
- threshold: "ema_params": {"decay": 0.9999, "decay_type": "threshold"}
- exp: "ema_params": {"decay": 0.9999, "decay_type": "exp", "beta": 15}Adding your own decay schedule
It is possible to bring your own decay schedule in SuperGradients. By subclassing from
IDecayFunction one can implement a custom
function:pyfrom super_gradients.training.utils.ema_decay_schedules import IDecayFunction, EMA_DECAY_FUNCTIONSclass LinearDecay(IDecayFunction):
def __init__(self, kwargs):
pass
def __call__(self, decay: float, step: int, total_steps: int) -> float:
"""
Compute EMA for specific training step following linear scaling rule [0..decay)
:param decay: The maximum decay value.
:param step: Current training step. The unit-range training percentage can be obtained by
step / total_steps.
:param total_steps: Total number of training steps.
:return: Computed decay value for a given step.
"""
training_progress = step / total_steps
return decay * training_progressEMA_DECAY_FUNCTIONS["linear"] = LinearDecay
How EMA weights are saved
When EMA is enabled, saved checkpoints will contain additional
ema_net attribute.
Weights for EMA model are saved under ema_net.
A regular (non-averaged) model weights are saved as net key in checkpoint as usual.When instantiating a model via
models.get, a function will check whether ema_net is present in the checkpoint.
In case ema_net is in checkpoint, the model will be initialized using EMA weights; otherwise a model will load initialized from regular weights saved in net.
Knowledge Distillation
EMA is also supported in knowledge distillation. To enable it one should add following parameters to the
training_params similar to the above example:pyfrom super_gradients import KDTrainer
trainer = KDTrainer(...)
trainer.train(
training_params={"ema": True, "ema_params": {"decay": 0.9999, "decay_type": "constant"}, ...},
...
)
---
Documentation/Source/Example Classification (documentation/source/Example_Classification.md)
Training a Classification Model and Transfer Learning
In this example we will use SuperGradients to train from scratch a ResNet18 model on the CIFAR10 image classification
dataset. We will also fine-tune the same model via transfer learning with weights pre-trained on the ImageNet dataset.
Quick installation
For this example, the only necessary package is super-gradients. Installing super-gradients will also
install all dependencies required to run the code in this example.
textpip install super-gradients
1. Experiment setup
First, we will initialize the
Trainer. It handles:
- Model training
- Evaluating test data
- Making predictions
- Saving and managing checkpoints
To initialize it, you need:
- Experiment Name: A unique identifier for your training experiment.
- Checkpoint Root Directory (
ckpt_root_dir): The directory where checkpoints, logs, and tensorboards are saved. While optional, if unspecified, it assumes the presence of a 'checkpoints' directory in your project's root.pythonfrom super_gradients import Trainerexperiment_name = "resnet18_cifar10_example"
CHECKPOINT_DIR = '/path/to/checkpoints/root/dir'
trainer = Trainer(experiment_name=experiment_name, ckpt_root_dir=CHECKPOINT_DIR)
2. Understanding the Checkpoint Structure
Checkpoints are crucial for progressive training, debugging, and model deployment. SuperGradients organizes them in a structured manner. Here's what the directory hierarchy looks like under your specified
ckpt_root_dir:text<ckpt_root_dir>
β
βββ <experiment_name>
β β
β ββββ <run_dir>
β β ββ ckpt_best.pth # Best performance during validation
β β ββ ckpt_latest.pth # End of the most recent epoch
β β ββ average_model.pth # Averaged over specified epochs
β β ββ ckpt_epoch_*.pth # Checkpoints from specific epochs (like epoch 10, 15, etc.)
β β ββ events.out.tfevents.* # Tensorflow run artifacts
β β ββ log_<timestamp>.txt # Trainer logs of the specific run
β β
β ββββ <other_run_dir>
β ββ ...
β
ββββ <other_experiment_name>
β
ββββ <run_dir>
β ββ ...
β
ββββ <another_run_dir>
ββ ...
In this structure:
-
ckpt_best.pth: Saved whenever there's an improvement in the specified validation metric.
- ckpt_latest.pth: Updated at the end of every epoch.
- average_model.pth: Averaged checkpoint, created if average_best_models parameter is set to True.For more information, check out the dedicated page.
2. Dataset and dataloaders
The dataset used in this example is the CIFAR10 image classification dataset.
SuperGradients provides a pool of standard datasets and dataloaders readily available for quick and easy usage.
SuperGradients also downloads the datasets when necessary, and gracefully handles the creation of the dataloaders,
with a pre-made training recipe specifically tailored for the dataset and model architecture.
<ins>Note:</ins> The SuperGradients trainer is compatible with PyTorch dataloaders and dataset objects. While it
is outside the scope of this example, it is worth remembering that custom dataloaders and datasets
can be employed when necessary.
2.A. Default dataloader from SuperGradients
As can be seen in the code snippet below, creating the training and validation dataloaders using SuperGradients'
default implementation is as easy as writing two lines of code:
textfrom super_gradients.training import dataloaderstrain_dataloader = dataloaders.get(name="cifar10_train", dataset_params={}, dataloader_params={"num_workers": 2})
valid_dataloader = dataloaders.get(name="cifar10_val", dataset_params={}, dataloader_params={"num_workers": 2})
Here, we call the
get() function twice, for the training and validation dataloaders. The function's parameters are:*
name - a string representing the name of the desired dataloader, out of a variety of different pre-made dataloaders
provided by SuperGradients. In this example, we use the pre-made CIFAR10 training and validation dataloaders.
* dataset_params - a dictionary of dataset-related parameters. Used to override the default parameters defined in the
training recipe. Later in this example we will show how this can be used to change the transforms
applied to the images.
* dataloader_params - a dictionary of dataloader-related parameters. Used to override the default parameters defined
in the training recipe. Here, as an example, we set the number of workers to 2.
* dataset - a torch.utils.data.Dataset object. Used when employing a custom dataset implementation. This parameter
cannot be passed together with the name or dataset_params parameters.We can always print the parameter values of the dataloader and its related dataset:
textimport pprintprint("Dataloader parameters:")
pprint.pprint(train_dataloader.dataloader_params)
print("Dataset parameters:")
pprint.pprint(train_dataloader.dataset.dataset_params)
Expected output:
textDataloader parameters:
{
"batch_size": 256,
"drop_last": False,
"num_workers": 2,
"pin_memory": True,
"shuffle": True
}Dataset parameters:
{
"download": True,
"root": "./data/cifar10",
"target_transform": None,
"train": True,
"transforms": [
{"RandomCrop": {"size": 32, "padding": 4}},
"RandomHorizontalFlip",
"ToTensor",
{"Normalize": {"mean": [0.4914, 0.4822, 0.4465], "std": [0.2023, 0.1994, 0.201]}},
]
}
When the
get() function is called as above, SuperGradients will attempt to download the CIFAR10 dataset for us. We
can expect to see an output as follows:<img src="images/classification_example_dataset_download.png" width="750">
After the dataloaders are defined, we can iterate them to extract batches of images and their corresponding labels.
This is useful for several purposes, such as visualization, verifying tensor shapes, and more. For example, visualization:
pythonfrom matplotlib import pyplot as pltdef show(images, labels, classes, rows=6, columns=5):
fig = plt.figure(figsize=(10, 10))
for i in range(1, columns * rows + 1):
fig.add_subplot(rows, columns, i)
plt.imshow(images[i-1].permute(1, 2, 0).clamp(0, 1))
plt.xticks([])
plt.yticks([])
plt.title(f"{classes[labels[i-1]]}")
plt.show()
images_train, labels_train = next(iter(train_dataloader))
show(images_train, labels_train, classes=train_dataloader.dataset.classes)
Output:
<img src="images/classification_example_data_vis.png" width="500">
As can be seen, the images are normalized. The normalization process is defined, among other things, as part of the
default training recipe SuperGradients uses for the CIFAR10 dataset. As we will see in the following section,
SuperGradients makes it a trivial task to override all, or part, of the different dataset and dataloader
parameters, allowing for control over the flexibility vs. ease-of-use tradeoff.
For completion of this section, let's print the tensors' shapes:
pythonprint(f'Training image tensor shape: {images_train.shape}')
print(f'Training labels tensor shape: {labels_train.shape}')
output:
textTraining image tensor shape: torch.Size([256, 3, 32, 32])
Training labels tensor shape: torch.Size([256])
As we can see, the default batch size for the training dataloader is 256.
2.B. Override parameters in dataset and dataloaders creation
To showcase the flexibility SuperGradients allows in customizing the different trainer
components, we will override the list of transforms that are used in the dataset. To define a list of
transformations to apply, we will use
torchvision's transforms. This also serves to show
the seamless integration SuperGradients allows with different PyTorch components. For the sake of
visualization, the only transform we apply is ToTensor(), which simply converts the input
images into PyTorch tensors.pythonfrom torchvision import transforms as Ttransforms_list = [T.ToTensor()]
vis_dataloader = dataloaders.get("cifar10_train",
dataset_params={"transforms": transforms_list},
dataloader_params={"num_workers": 2})
images, labels = next(iter(vis_dataloader))
show(images, labels, classes=train_dataloader.dataset.classes)
Notice that the only difference in the dataloader's definition is that here, the
dataset_params parameter
is passed as a dictionary defining the parameters to override. The result of running the above code:<img src="images/classification_example_transforms_data_vis.png" width="500">
The effect of changing the transforms can be seen in the images - now, without normalization, the objects are more
observable in the images.
3. Architecture definition
In this example, we train the model with the ResNet18 architecture.
SuperGradients provides out-of-the-box implementations of many architectures for classification tasks. With just one
line of code we can define a model with the chosen architecture. A list of all available architectures can be found
here.
pythonfrom super_gradients.training import models
from super_gradients.common.object_names import Modelsmodel = models.get(model_name=Models.RESNET18, num_classes=10)
Notice that, similar to obtaining a pre-defined dataloader, here we use
super_gradients.training.models's get()
function. In the above code, two parameters are passed to the function:*
model_name - A string defining the model's architecture name, out of the list of architectures SuperGradients
provides.
* num_classes - An integer representing the number of classes the model should learn to predict. Affects the
architecture's structure.Some additional parameters the
get() function supports:*
arch_params - A dictionary used to override the default architecture parameters, such as the number of residual
blocks.
* checkpoint_path - A string defining the path to an external checkpoint to be loaded. Can be absolute or relative.
If provided, will automatically attempt to load the checkpoint.
* pretrained_weights - A string defining the name of a dataset on which the model was pre-trained on, for
fine-tuning and transfer learning. The pretrained_weights and checkpoint_path parameters
are mutually exclusive.For more available parameters, refer to the function's docstring.
In this example, we have defined the model with one of SuperGradient's readily available architectures. As was
already noted in previous sections, SuperGradients is highly compatible with PyTorch. Defining the model's architecture
is not an exception - we can seamlessly use a custom architecture, i.e. a
torch.nn.Module object, for maximum
flexibility, although it is out of the scope of this example.4. Training setup
We have defined the trainer, datasets, dataloaders, and model architecture. Before we can start training, we need to
define the training parameters. As with the other parameters, SuperGradients provides training parameters optimized for
this use-case. For more recommended training parameters you can have a look at our recipes
here.
Obtaining the training parameters is as easy as writing a single line of code:
pythonfrom super_gradients.training import training_hyperparamstraining_params = training_hyperparams.get(config_name="training_hyperparams/cifar10_resnet_train_params")
We notice the repeatability in the code usage - to obtain the training parameters, we again call
the
get() function. This function accepts two parameters:*
config_name - A string defining the .yaml config filename in the recipes' directory.
* overriding_params - An optional parameter, a dictionary used to override the loaded training parameters.We can print the training parameters to see the different options:
pythonpprint.pprint("Training parameters:")
pprint.pprint(training_params)
Output (Training parameters):
text/ Detailed source-code truncated for AI context efficiency. /
As can be seen in the above output, there are numerous options to modify the training parameters
to affect the training process.
It is also possible to change training parameters after obtaining them, for example:
pythontraining_params["max_epochs"] = 15
training_params["sg_logger_params"]["launch_tensorboard"] = True
5. Training, checkpointing, and transfer learning
5.A. Training the model
We are all set to start training our model. Simply plug in the model, training and validation dataloaders,
and training parameters into the trainer's
train() function:pythontrainer.train(model=model,
training_params=training_params,
train_loader=train_dataloader,
valid_loader=valid_dataloader)
The training progress will be printed to the screen:
text[2023-02-01 20:57:27] INFO - sg_trainer_utils.py - TRAINING PARAMETERS:
- Mode: Single GPU
- Number of GPUs: 1 (4 available on the machine)
- Dataset size: 50000 (len(train_set))
- Batch size per GPU: 256 (batch_size)
- Batch Accumulate: 1 (batch_accumulate)
- Total batch size: 256 (num_gpus * batch_size)
- Effective Batch size: 256 (num_gpus batch_size batch_accumulate)
- Iterations per epoch: 195 (len(train_set) / total_batch_size)
- Gradient updates per epoch: 195 (len(train_set) / effective_batch_size)[2023-02-01 20:57:27] INFO - sg_trainer.py - Started training for 15 epochs (0/14)
Train epoch 0: 100%|ββββββββββ| 196/196 [00:18<00:00, 10.51it/s, Accuracy=0.262, CrossEntropyLoss=2.37, Top5=0.787, gpu_mem=0.371]
Validation epoch 0: 100%|ββββββββββ| 20/20 [00:03<00:00, 6.12it/s]
===========================================================
SUMMARY OF EPOCH 0
βββ Training
β βββ Accuracy = 0.262
β βββ CrossEntropyLoss = 2.3702
β βββ Top5 = 0.787
βββ Validation
βββ Accuracy = 0.3459
βββ CrossEntropyLoss = 1.8811
βββ Top5 = 0.871
===========================================================
At the beginning of the training, a summary of the training parameters is printed, where we can see
the training mode (CPU/single GPU/distributed training), the number of GPUs used, the training dataset
size, and more. The progress of each epoch's training and validation is displayed, along with the tracked metrics
(defined as part of the training recipe): accuracy, loss value, top5 error, and GPU memory consumption.
At the end of each epoch, a summary of the training and validation metrics is displayed, and in later epochs,
a comparison with the previous epochs is provided:
text===========================================================
SUMMARY OF EPOCH 15
βββ Training
β βββ Accuracy = 0.7594
β β βββ Best until now = 0.7458 (β 0.0136)
β β βββ Epoch N-1 = 0.7458 (β 0.0136)
β βββ CrossEntropyLoss = 0.686
β β βββ Best until now = 0.7187 (β -0.0327)
β β βββ Epoch N-1 = 0.7187 (β -0.0327)
β βββ Top5 = 0.9867
β βββ Best until now = 0.9849 (β 0.0019)
β βββ Epoch N-1 = 0.9849 (β 0.0019)
βββ Validation
βββ Accuracy = 0.7425
β βββ Best until now = 0.746 (β -0.0035)
β βββ Epoch N-1 = 0.7306 (β 0.0119)
βββ CrossEntropyLoss = 0.7331
β βββ Best until now = 0.7315 (β 0.0016)
β βββ Epoch N-1 = 0.8048 (β -0.0717)
βββ Top5 = 0.9831
βββ Best until now = 0.9838 (β -0.0007)
βββ Epoch N-1 = 0.9818 (β 0.0013)===========================================================
At the end of each epoch, the different logs and checkpoints are saved in the path defined by
ckpt_root_dir and
experiment_name.In the epoch summary shown above, we can see that the validation accuracy is 73%, which is
not very high. To get better insights as to what is happening, we turn to the tensorboard logs.
5.B. Tensorboard logs
To view the experiment's tensorboard logs, type the following command in the terminal from the
experiment's path:
bashtensorboard --logdir='.'
(Alternatively, run the command from anywhere with the experiment's full path).
SuperGradients logs many useful metrics to tensorboard, including CPU and GPU usage, learning rate
scheduling, training and validation losses and other metrics, and many more. For the purpose of this example,
let us examine the training and validation loss:
<img src="images/classification_example_train_loss_initial.png" width="500">
<img src="images/classification_example_valid_loss_initial.png" width="500">
As can be seen in the graphs, the training (and validation) loss did not converge before training ended. This means
that training the model for additional epochs will probably improve its performance. Earlier, when modifying the
training parameters, we set
max_epochs = 15. Let us continue training the model for an additional 10 epochs.5.C. Continue training from a checkpoint
To continue training from a checkpoint, we utilize the
models.get() function's checkpoint_path parameter.
The provided checkpoint path includes the checkpoint file we wish to load. In this example, since we want
to continue from the last checkpoint, we will load the ckpt_latest.pth checkpoint. Additionally, we want to let
the trainer know that we are continuing training and not starting from the first epoch. This is done
by setting the resume training parameter to True. Finally, we set the new max_epochs training
parameter, and train the model once more.pythonimport osmodel = models.get(model_name=Models.RESNET18,
num_classes=10,
checkpoint_path=os.path.join(CHECKPOINT_DIR, experiment_name, 'ckpt_latest.pth'))
training_params["resume"] = True
training_params["max_epochs"] = 25
trainer.train(model=model,
training_params=training_params,
train_loader=train_dataloader,
valid_loader=valid_dataloader)
We can see that the training continues for 10 epochs, resuming from epoch 15:
text[2023-02-01 21:21:16] INFO - sg_trainer_utils.py - TRAINING PARAMETERS:
- Mode: Single GPU
- Number of GPUs: 1 (4 available on the machine)
- Dataset size: 50000 (len(train_set))
- Batch size per GPU: 256 (batch_size)
- Batch Accumulate: 1 (batch_accumulate)
- Total batch size: 256 (num_gpus * batch_size)
- Effective Batch size: 256 (num_gpus batch_size batch_accumulate)
- Iterations per epoch: 195 (len(train_set) / total_batch_size)
- Gradient updates per epoch: 195 (len(train_set) / effective_batch_size)[2023-02-01 21:21:16] INFO - sg_trainer.py - Started training for 10 epochs (15/24)
Train epoch 15: 100%|ββββββββββ| 196/196 [00:18<00:00, 10.52it/s, Accuracy=0.764, CrossEntropyLoss=0.668, Top5=0.987, gpu_mem=0.422]
Validation epoch 15: 100%|ββββββββββ| 20/20 [00:03<00:00, 6.10it/s]
===========================================================
SUMMARY OF EPOCH 15
βββ Training
β βββ Accuracy = 0.7644
β βββ CrossEntropyLoss = 0.6684
β βββ Top5 = 0.9865
βββ Validation
βββ Accuracy = 0.7539
βββ CrossEntropyLoss = 0.7271
βββ Top5 = 0.9841
===========================================================
Finally, the model stops training after completing 25 epochs:
text===========================================================
SUMMARY OF EPOCH 25
βββ Training
β βββ Accuracy = 0.8177
β β βββ Best until now = 0.8147 (β 0.003)
β β βββ Epoch N-1 = 0.8147 (β 0.003)
β βββ CrossEntropyLoss = 0.5211
β β βββ Best until now = 0.5281 (β -0.007)
β β βββ Epoch N-1 = 0.5281 (β -0.007)
β βββ Top5 = 0.9921
β βββ Best until now = 0.9919 (β 0.0002)
β βββ Epoch N-1 = 0.9919 (β 0.0002)
βββ Validation
βββ Accuracy = 0.8201
β βββ Best until now = 0.7873 (β 0.0328)
β βββ Epoch N-1 = 0.7534 (β 0.0667)
βββ CrossEntropyLoss = 0.525
β βββ Best until now = 0.6145 (β -0.0895)
β βββ Epoch N-1 = 0.7517 (β -0.2266)
βββ Top5 = 0.9907
βββ Best until now = 0.9883 (β 0.0024)
βββ Epoch N-1 = 0.983 (β 0.0077)===========================================================
We can see that the validation accuracy is now 82%. Much better.
5.C. Transfer learning
So far, we trained a model from scratch. More formally, the model's weights were randomly initialized. For easy tasks
and large datasets this is usually sufficient. In other cases, especially when not a lot of data is available for
training, we would like to take advantage of knowledge gained from other sources. This is called transfer learning,
and it has many forms and variations. In this example we will take a look at the simplest form of transfer learning:
fine-tuning a model initialized with pre-trained weights. Specifically, we will initialize our model
with weights pre-trained on the ImageNet dataset.
SuperGradients provides a variety of pre-trained weights readily available for fine-tuning different models. To
initialize our model with pre-trained weights provided by SuperGradients, only a small change to the existing code
is needed:
textmodel = models.get(model_name=Models.RESNET18, num_classes=10, pretrained_weights="imagenet")
In the above code, we provided the
pretrained_weights parameter to the models.get() function. This parameter accepts
a string representing the name of the dataset that the weights were pre-trained on. Note that this
parameter is mutually exclusive with the checkpoint_path parameter.The rest of the training pipeline is the same as above. For comparison with the previous model, we will train this model
for 25 epochs also. Let us look how the model performed:
text===========================================================
SUMMARY OF EPOCH 25
βββ Training
β βββ Accuracy = 0.8242
β β βββ Best until now = 0.8267 (β -0.0025)
β β βββ Epoch N-1 = 0.8267 (β -0.0025)
β βββ CrossEntropyLoss = 0.5035
β β βββ Best until now = 0.4998 (β 0.0037)
β β βββ Epoch N-1 = 0.4998 (β 0.0037)
β βββ Top5 = 0.9924
β βββ Best until now = 0.9917 (β 0.0007)
β βββ Epoch N-1 = 0.9917 (β 0.0007)
βββ Validation
βββ Accuracy = 0.8377
β βββ Best until now = 0.8062 (β 0.0315)
β βββ Epoch N-1 = 0.806 (β 0.0317)
βββ CrossEntropyLoss = 0.4834
β βββ Best until now = 0.5731 (β -0.0897)
β βββ Epoch N-1 = 0.5785 (β -0.0952)
βββ Top5 = 0.9924
βββ Best until now = 0.9903 (β 0.0021)
βββ Epoch N-1 = 0.9903 (β 0.0021)===========================================================
As we can see, the validation accuracy improved by 1.7% compared to the randomly initialized model.
To achieve a greater improvement with pre-trained weights, sometimes careful tuning of the training hyperparameters
is required.
6. Predictions with the trained model
Now that we have a trained model with reasonable performance, we can use it to make predictions on new data.
First, let's import some packages:
pythonfrom PIL import Image
import torch
import numpy as np
import requests
Next, we load the model with the trained weights, and put it into evaluation mode. Notice that now we load the
ckpt_best.pth checkpoint.pythonmodel = models.get(model_name=Models.RESNET18,
num_classes=10,
checkpoint_path=os.path.join(CHECKPOINT_DIR, experiment_name, 'ckpt_best.pth'))
model.eval()
We want to test the model on an image of one the classes the model was trained on. As an example, let us test how the
model handles an image of a frog, taken from the Aquarium of the Pacific website.
The loaded image must undergo the same transformations as the training images for the model to work well:
pythonurl = "https://www.aquariumofpacific.org/images/exhibits/Magnificent_Tree_Frog_900.jpg"
image = np.array(Image.open(requests.get(url, stream=True).raw))transforms = T.Compose([
T.ToTensor(),
T.Normalize(mean=(0.4914, 0.4822, 0.4465), std=(0.2023, 0.1994, 0.2010)),
T.Resize((32, 32))
])
input_tensor = transforms(image).unsqueeze(0).to(next(model.parameters()).device)
Next, to obtain the model's predictions we simply run the following line of code:
pythonpredictions = model(input_tensor)
Let's see what the model predicted:
pythonplt.xlabel(train_dataloader.dataset.classes[torch.argmax(predictions)])
plt.imshow(image)
plt.show()
<img src="images/classification_example_frog_prediction.png" width="500">
As we can see, the model correctly predicted that the input image is an image of a frog.
7. Complete code
For completion of this example, we provide a complete working code for training, continuing training from
a saved checkpoint, and predicting with the trained model. Simply change the
CHECKPOINT_DIR variable and run the
script:text/ Detailed source-code truncated for AI context efficiency. /
---
Documentation/Source/Example Training An External Model (documentation/source/Example_Training-an-external-model.md)
Training an external model
In this example we will use SuperGradients to train a deep learning segmentation model to extract human portraits from
images, i.e., to remove the background from the image.
We will show how SuperGradients allows seamless integration of
an external model, dataset, loss function, and metric into the training pipeline.
Quick installation
For this example, the only necessary package is super-gradients. Installing super-gradients will also
install all dependencies required to run the code in this example.
bashpip install super-gradients
1. Dataset
The dataset we will use in this example is the
AISegment dataset, available to download for free from
Kaggle.
The dataset contains 34,427 RGB images of human portraits and their corresponding soft masks.
The provided images are center-cropped to a unified shape of 600x800.
1.A. Data preparation
The original structure of the data on the disk is as follows:
textdata
ββββaisegment-matting
β ββββmatting
β β ββββ1803290511
β β β ββββmatting_00000000
β β β β 1803290511-00000459.png
β β β β 1803290511-00000458.png
β β β β ..
β β ββββ1803290444
β β ..
β ββββclip_img
β β ββββ1803290511
β β β ββββclip_00000000
β β β β 1803290511-00000459.jpg
β β β β 1803290511-00000458.jpg
β β β β ..
β β ββββ1803290444
β β ..
In the structure shown above, the
clip_img directory contains the portrait images, divided into many
sub-directories. The matting directory is structured similar to the clip_img directory, containing
the corresponding masks. This structure is not particularly convenient for data loading, therefore
we will first rearrange the data with the following script, which only requires changing the data_path and
out_path variables:text/ Detailed source-code truncated for AI context efficiency. /
The original data may have images without their corresponding masks. Therefore, the above code first finds all existing
image-mask pairs. It then splits the data into train/validation/test sets, where the size of each of the validation and
test sets is 10% of the total number of samples. The samples in each of the sets are then copied to the output path,
structured as follows:
textdata
ββββtrain
β 1803290511-00000029.jpg
β 1803290511-00000029.png
β ..
ββββval
β 1803290444-00000292.jpg
β 1803290444-00000292.png
β ..
ββββtest
β 1803290443-00000294.jpg
β 1803290443-00000294.png
β ..
Each of the train/val/test directories contain all image (.jpg) files and their corresponding mask (.png) files.
1.B. PyTorch Dataset
In some cases, we may want to have full control over the process of loading and pre-processing the training data.
SuperGradients is fully compatible with PyTorch data loaders, which allows for seamless integration of custom dataset
implementations for maximum flexibility.
We will first present the complete Dataset class implementation, and then break it down to fully understand what is
going on.
text/ Detailed source-code truncated for AI context efficiency. /
First, we can see that our dataset class inherits from the
torch.utils.data.Dataset class. To initialize the dataset
object, two parameters must be provided:*
data_path - the full path to the data's root directory
* split - a string indicating whether this is the 'train', 'val', or 'test' split of the dataAdditional parameters include
input_height and input_width, the fixed height and width, respectively, for resizing
the input images and masks. These are set to 256 by default. At the end of the initialization function, the image and
mask path pairs are determined according to the split.Next, let's see what happens when we retrieve an item from the dataset via the
__getitem__() function.
First, an image and its corresponding mask are loaded according to the idx parameter:pythonimage = Image.open(self.path_pairs[idx][0]).convert('RGB')
mask = Image.open(self.path_pairs[idx][1]).split()[-1]
Notice that we take only the mask's last channel. The mask's original color format is RGBA, and the alpha channel should be
used as the soft mask for segmentation according to the
dataset's description.
Next, we apply transformations to the image and mask according to the split:
pythonif self.split == 'train':
seg_transforms = torch_transforms.Compose([
SegRandomFlip(prob=0.5),
SegColorJitter(brightness=0.5, contrast=0.5, saturation=0.5),
SegResize(h=self.input_height, w=self.input_width)
])
else:
seg_transforms = SegResize(h=self.input_height, w=self.input_width)transformed_pair = seg_transforms({"image": image, "mask": mask})
image, mask = transformed_pair['image'], transformed_pair['mask']
On training images and masks we apply data augmentation: a random horizontal flip with probability 0.5, and
random color jitter which randomly changes the image's brightness, contrast, and saturation. Both the training and the
other splits' images and masks undergo resizing according to
input_height and input_width.The transforms in the above code are all SuperGradients transforms, which are built upon PyTorch's
torchvision.transforms. This allows for maximum compatibility with PyTorch components. For example, we can see that
the transforms are sequentially composed using torchvision's Compose class. Notice also that the transform names are
all prefixed with Seg - meaning that they are specifically designed for segmentation data. These transforms take care
to apply the same transformation to the image and the mask when required, or only apply the transformation to the image
otherwise. For example, if SegRandomFlip() flips the image, the mask will be flipped as well. SegColorJitter() only
transforms the image. SegResize() resizes both the image and the mask.Next, we convert the image and the mask to PyTorch Tensors, and normalize the image using the
NORMALIZATION_MEANS and
NORMALIZATION_STDS:pythonimage_transform = torch_transforms.Compose([
torch_transforms.ToTensor(),
torch_transforms.Normalize(self.NORMALIZATION_MEANS, self.NORMALIZATION_STDS)
])
mask_transform = torch_transforms.ToTensor()image_tensor, mask_tensor = image_transform(image), mask_transform(mask)
These transforms are applied regardless of the current dataset split. Notice that here we use torchvision transforms.
This serves to show the high degree of flexibility SuperGradients allows, as its transforms are based on PyTorch's
torchvision transforms which may be used interchangeably.
1.C. Data visualization
To conclude this section, let's visualize some images and their masks to test our AISegmentDataset implementation.
First, we instantiate two dataset objects, for the training and validation splits, and extract the first sample from
each:
pythonimport matplotlib.pyplot as pltdata_path = '/path/to/arranged/data/dir'
train_dataset = AISegmentDataset(data_path=data_path, split='train')
val_dataset = AISegmentDataset(data_path=data_path, split='val')
train_image, train_mask = train_dataset[0]
val_image, val_mask = val_dataset[0]
Let's first visualize the validation image and mask:
pythonfigure = plt.figure()
figure.add_subplot(1, 2, 1)
plt.title("Image")
plt.axis("off")
plt.imshow(val_image.permute(1, 2, 0))
figure.add_subplot(1, 2, 2)
plt.title("Mask")
plt.axis("off")
plt.imshow(val_mask.squeeze(), cmap='gray')plt.show()
<img src="images/external_model_example_val_vis.png" width="750">
And the training image and mask:
pythonfigure = plt.figure()
figure.add_subplot(1, 2, 1)
plt.title("Image")
plt.axis("off")
plt.imshow(train_image.permute(1, 2, 0))
figure.add_subplot(1, 2, 2)
plt.title("Mask")
plt.axis("off")
plt.imshow(train_mask.squeeze(), cmap='gray')plt.show()
<img src="images/external_model_example_train_vis.png" width="750">
We can see the effect of the color jitter transform on the image.
2. Model architecture
For this example we will employ the U-Net architecture. U-Net and its variants are a
popular choice for many image segmentation tasks. It is a fully-convolutional architecture,
consisting of an expanding and a contracting path with skip connections between the encoder
and decoder blocks. In this example we will demonstrate how we can easily integrate an external PyTorch model
as part of SuperGradients' training pipeline. To this end, we will use a
U-Net implementation loaded directly from
PyTorch Hub:
pythonmodel = torch.hub.load('mateuszbuda/brain-segmentation-pytorch', 'unet', in_channels=3, out_channels=1,
init_features=32, pretrained=False)
Since our model's inputs are RGB images, we set
in_channels=3. This is a binary segmentation task, therefore
out_channels=1. The init_features parameter determines the number of kernels in the first block's convolution
layers. The number of kernels is doubled in each consecutive encoder block. We also set pretrained=False since
in this case we do not want to use pre-trained weights.Let's check our model's type:
pythonprint(type(model))
text<class 'unet.UNet'>
SuperGradients is compatible with models of type
torch.nn.Module. Just to be sure, let's verify that our
model is of the correct type:pythonprint(type(model).__bases__)
text(<class 'torch.nn.modules.module.Module'>,)
Finally, let's print the model to see its components:
pythonprint(model)
text/ Detailed source-code truncated for AI context efficiency. /
3. Loss function
While a popular choice for binary segmentation loss function is the binary cross-entropy (BCE) loss, often it is
the intersection-over-union (IoU), or the Jaccard Index, that serves
as a measure of success. In this example, we will use a combination of the BCE and IoU as the loss function. While
SuperGradients provides, among many other losses, an implementation of the combined BCE and Dice loss, which could be
used for our purposes as well, we will show how we can define our own user-defined loss function and train our model
with it using SuperGradients.
Similar to using an external model, the custom loss function's class must inherit from
torch.nn.Module. The
forward() function's first parameter needs to be the predictions tensor and the second parameter needs to be the
target tensor.pythonimport torch
import torch.nn as nn
class CustomIoU(torch.nn.Module):
def __init__(self):
super(CustomIoU, self).__init__()
def forward(self, preds, target):
intersection = torch.sum(target * preds)
union = torch.sum(target) + torch.sum(preds) - intersection + 1e-5
iou = intersection / union
return iou
class CustomSegLoss(torch.nn.Module):
def __init__(self, bce_weight=1, iou_weight=1):
super(CustomSegLoss, self).__init__()
self.bce_weight = bce_weight
self.iou_weight = iou_weight
self.bce_loss = nn.BCELoss()
self.iou_func = CustomIoU()
def forward(self, preds, target):
bce_loss = self.bce_loss(preds, target)
iou_loss = 1.0 - self.iou_func(torch.gt(preds, 0.5).long(), torch.gt(target, 0.5).long())
return self.bce_weightbce_loss + self.iou_weightiou_loss
Notice that here the BCE loss term is obtained simply by using PyTorch's
BCELoss(). To compute the IoU score, we
have implemented an auxiliary class CustomIoU, which implements a naive, differentiable IoU function. Note that
in binary segmentation tasks, we are usually interested only in the foreground IoU. Therefore, CustomIoU disregards
the background IoU. To compute the IoU, the pixel values in both images should be binary, i.e., 0's and 1's.
Since in the model's forward function a sigmoid function is already applied to the output, we only need to binarize
both the predictions and the target tensors with a threshold of 0.5 (remember, we are using soft masks). When
measuring segmentation performance, higher IoU is better. Since IoU score is a number in [0, 1], we simply compute
IoU loss = 1 - IoU to make it a valid loss function for gradient-descent optimization.The overall loss is a weighted sum of the BCE and the IoU loss terms, with weights
bce_weight and iou_weight,
respectively. With just a few lines of code we have defined our own custom loss function. Although here we could have used a similar
loss function provided by SuperGradients, in many other cases a more complex and specific loss function is desired.
This example serves to show how we may define any loss function to fit our needs and seamlessly integrate it with the
SuperGradients training pipeline.
4. Custom IoU metric
Since our measure of success is the IoU, we would like to tell SuperGradients to log and track this metric during
training. This metric would also be used to determine the best model at the end of every epoch for checkpointing.
We note that SuperGradients provides many built-in metrics, including variants of the IoU. However, in this example
we aim to show the ease at which we can incorporate external metrics into our pipeline. SuperGradients supports any
metric of type
torchmetrics.Metric.The metric we will use is torchmetrics'
JaccardIndex. However, recall that we use soft masks.
JaccardIndex requires the target
tensor's elements to be integers. Also, as noted in the previous section, since this is a binary segmentation
task, we are interested in the foreground IoU. Therefore, we will need to modify the metric a bit. For more details
about implementing a custom metric using torchmetrics, see
here.pythonclass SoftIoU(JaccardIndex):
def __init__(self, kwargs):
super().__init__(reduction='none', kwargs) def update(self, preds: torch.Tensor, target: torch.Tensor):
target = torch.gt(target, 0.5).long()
super().update(preds, target)
def compute(self):
return super().compute()[1]
We have defined our
SoftIoU class which inherits from JaccardIndex. The only modifications we introduced to
JaccardIndex are:1. The target tensor is binarized with a threshold of 0.5:
target = torch.gt(target, 0.5).long()
2. To get the IoU of the foreground alone, we set reduction='none' in the __init__ function. This means that
instead of computing the mean of the background and foreground IoUs, both values are returned, and in the compute()
function we only take the second element, which corresponds to the foreground.Our custom metric is now ready to use with our training pipeline.
5. Experiment configuration
Trainer
First, we will initialize the Trainer. It handles:
- Model training
- Evaluating test data
- Making predictions
- Saving and managing checkpoints
To initialize it, you need:
- Experiment Name: A unique identifier for your training experiment.
- Checkpoint Root Directory (
ckpt_root_dir): The directory where checkpoints, logs, and tensorboards are saved. While optional, if unspecified, it assumes the presence of a 'checkpoints' directory in your project's root.pythonfrom super_gradients import Trainerexperiment_name = "aisegment_example"
CHECKPOINT_DIR = '/path/to/checkpoints/root/dir'
trainer = Trainer(experiment_name=experiment_name, ckpt_root_dir=CHECKPOINT_DIR)
Understanding the Checkpoint Structure
Checkpoints are crucial for progressive training, debugging, and model deployment. SuperGradients organizes them in a structured manner. Here's what the directory hierarchy looks like under your specified
ckpt_root_dir:text<ckpt_root_dir>
β
βββ <experiment_name>
β β
β ββββ <run_dir>
β β ββ ckpt_best.pth # Best performance during validation
β β ββ ckpt_latest.pth # End of the most recent epoch
β β ββ average_model.pth # Averaged over specified epochs
β β ββ ckpt_epoch_*.pth # Checkpoints from specific epochs (like epoch 10, 15, etc.)
β β ββ events.out.tfevents.* # Tensorflow run artifacts
β β ββ log_<timestamp>.txt # Trainer logs of the specific run
β β
β ββββ <other_run_dir>
β ββ ...
β
ββββ <other_experiment_name>
β
ββββ <run_dir>
β ββ ...
β
ββββ <another_run_dir>
ββ ...
In this structure:
-
ckpt_best.pth: Saved whenever there's an improvement in the specified validation metric.
- ckpt_latest.pth: Updated at the end of every epoch.
- average_model.pth: Averaged checkpoint, created if average_best_models parameter is set to True.For more information, check out the dedicated page.
Dataloaders
Next, we initialize the PyTorch dataloaders for our datasets:
pythonfrom torch.utils.data import DataLoadertrain_dataloader = DataLoader(train_dataset, batch_size=16, shuffle=True, num_workers=2)
val_dataloader = DataLoader(val_dataset, batch_size=16, shuffle=False, num_workers=2)
Training Hyperparameters
And lastly, we need to define the training hyperparameters:
pythontrain_params = {
"max_epochs": 100,
"lr_mode": "CosineLRScheduler",
"initial_lr": 0.001,
"optimizer": "Adam",
"loss": CustomSegLoss(),
"metric_to_watch": "SoftIoU",
"greater_metric_to_watch_is_better": True,
"train_metrics_list": [SoftIoU(num_classes=2)],
"valid_metrics_list": [SoftIoU(num_classes=2)]
}
Notice that the training hyperparameters must be defined as a dictionary with the hyperparameter names as keys.
The dictionary defines the hyperparameters that we want to override. All other hyperparameters retain their default
values defined by SuperGradients. The list of all training hyperparameters and their default value
can be found here.
The
metric_to_watch hyperparameter defines the metric used to determine the best model at the end of every epoch.
We simply set it as a string representing the name of our custom metric SoftIoU. Greater IoU is better, therefore we
set greater_metric_to_watch_is_better=True. The IoU metric will also be logged and tracked during training and
validation, as determined by train_metrics_list and valid_metrics_list. The loss hyperparameter tells
SuperGradients which loss function to use during training. To use one of the many loss functions provided by
SuperGradients, we set this hyperparameter as a string representing the loss function's name. However, here we want
to use our custom loss function. Therefore, we simply set the hyperparameter as a CustomSegLoss() object. The above code shows the simplicity of integrating external, user-defined components into the SuperGradients training
pipeline. We simply plugged instantiations of our custom loss and metric into the hyperparameters dictionary,
and we are ready to go.
6. Training
6.A. Training the model
We are all set to start training our model. Simply plug in the model, training and validation dataloaders,
and training parameters into the trainer's
train() function:pythontrainer.train(model=model,
training_params=train_params,
train_loader=train_dataloader,
valid_loader=val_dataloader)
The training progress will be printed to the screen:
text[2023-02-06 11:44:35] INFO - sg_trainer.py - Started training for 100 epochs (0/99)Train epoch 0: 100%|ββββββββββ| 3443/3443 [19:16<00:00, 2.98it/s, CustomSegLoss=0.24, SoftIoU=0.9, gpu_mem=1.81]
Validation epoch 0: 100%|ββββββββββ| 431/431 [00:42<00:00, 10.23it/s]
===========================================================
SUMMARY OF EPOCH 0
βββ Training
β βββ Customsegloss = 0.2466
β βββ Softiou = 0.9
βββ Validation
βββ Customsegloss = 0.1367
βββ Softiou = 0.9483
===========================================================
The progress of each epoch's training and validation is displayed, along with the values of our custom metric and
loss function, and GPU memory consumption. At the end of each epoch, a summary of the training and validation metrics
is displayed, and in later epochs, a comparison with the previous epochs is provided:
text===========================================================
SUMMARY OF EPOCH 5
βββ Training
β βββ Customsegloss = 0.0915
β β βββ Best until now = 0.0945 (^[[32mβ -0.003^[[0m)
β β βββ Epoch N-1 = 0.0945 (^[[32mβ -0.003^[[0m)
β βββ Softiou = 0.9651
β βββ Best until now = 0.9 (β 0.0651^[[0m)
β βββ Epoch N-1 = 0.9639 (β 0.0011^[[0m)
βββ Validation
βββ Customsegloss = 0.0789
β βββ Best until now = 0.0924 (^[[32mβ -0.0135^[[0m)
β βββ Epoch N-1 = 0.0924 (^[[32mβ -0.0135^[[0m)
βββ Softiou = 0.9702
βββ Best until now = 0.9483 (β 0.0219^[[0m)
βββ Epoch N-1 = 0.9657 (β 0.0045^[[0m)===========================================================
At the end of each epoch, the different logs and checkpoints are saved in the path defined by
ckpt_root_dir and
experiment_name. Let's see how we can use Tensorboard to track training process.6.B. Tensorboard logs
To view the experiment's tensorboard logs, type the following command in the terminal from the
experiment's path:
bashtensorboard --logdir='.'
(Alternatively, run the command from anywhere with the experiment's full path).
SuperGradients logs many useful metrics to tensorboard, including CPU and GPU usage, learning rate
scheduling, training and validation losses and other metrics, and many more. For example, let's check how the training
process goes by looking at the training's custom loss value:
<img src="images/external_model_example_train_loss.png" width="700">
We can also check the validation set's IoU metric's value:
<img src="images/external_model_example_valid_iou.png" width="700">
7. Predictions with the trained model
Now that we have a trained model we can use it to make predictions on the test set. First, let's instantiate a test
dataset:
pythontest_dataset = AISegmentDataset(data_path=data_path, split='test')
By instantiating the test dataset, all required pre-processing is already handled for us. Let's choose a single sample
from the dataset:
pythonimage, mask = test_dataset[0]
image, mask = image.unsqueeze(0), mask.unsqueeze(0)
Notice that we have added the batch dimension to the tensors. Next, we set the model to evaluation mode, and obtain
the predicted mask. Since our model learned to predict soft masks, we apply a threshold of 0.5 to binarize the mask:
pythonmodel.eval()pred = model(image)
pred = torch.gt(pred, 0.5).long()
Finally, let's apply the predicted mask to the image:
pythonmasked_image = image*pred
Now let's visualize the image, mask, and masked image to see how our model performed:
pythonfigure = plt.figure()
figure.add_subplot(1, 3, 1)
plt.title("Image")
plt.axis("off")
plt.imshow(image.detach().squeeze(0).permute(1, 2, 0))
figure.add_subplot(1, 3, 2)
plt.title("Mask")
plt.axis("off")
plt.imshow(pred.detach().squeeze(0).permute(1, 2, 0), cmap='gray')
figure.add_subplot(1, 3, 3)
plt.title("Masked Image")
plt.axis("off")
plt.imshow(masked_image.detach().squeeze(0).permute(1, 2, 0))
plt.show()
<img src="images/external_model_example_prediction.png" width="700">
---
Documentation/Source/Experiment Management (documentation/source/experiment_management.md)
Experiment Management
Outline
1. Core Concepts
- Checkpoint Root Directory
- Experiments
- Runs
2. File Structure of Experiments
3. Utilities for Experiment Management
- Get the Absolute Path of a Run Directory
- Retrieve the Latest Run IDCore Concepts
Checkpoint Root Directory (
ckpt_root_dir)
- The main directory where all experiment outputs are housed.Experiments (
experiment_name)
- Symbolizes a distinct training recipe or configuration.
- Alter the experiment_name for transparency when updating your training recipe.
- Each training under the same experiment_name has its individual run directory, ensuring no overwrites.Runs (
run_id)
- Every individual training session is termed as a run.
- A unique run_id is generated for every training, regardless of identical parameters.
- Different trainings under the same experiment_name maintain distinct logs and checkpoints, courtesy of their separate run directories.File Structure of Experiments
text<ckpt_root_dir>
β
βββ <experiment_name>
β β
β ββββ <run_dir>
β β ββ ckpt_best.pth # Best performance during validation
β β ββ ckpt_latest.pth # End of the most recent epoch
β β ββ average_model.pth # Averaged over specified epochs
β β ββ ckpt_epoch_*.pth # Checkpoints from certain epochs (e.g., epoch 10, 15)
β β ββ events.out.tfevents.* # Tensorflow run artifacts
β β ββ log_<timestamp>.txt # Trainer logs of that particular run
β β
β ββββ <other_run_dir>
β ββ ...
β
ββββ <other_experiment_name>
β
ββββ <run_dir>
β ββ ...
β
ββββ <another_run_dir>
ββ ...
Utilities
#### A. Get the absolute path of a run directory
Manually navigate using
<ckpt_root_dir>/<experiment_name>/<run_dir> or utilize the following programmatic approach:
pythonfrom super_gradients.common.environment.checkpoints_dir_utils import get_checkpoints_dir_pathcheckpoints_dir_path = get_checkpoints_dir_path(experiment_name="<experiment_name>", run_id="<run_id>")
#### B. Get the latest run id
pythonfrom super_gradients.common.environment.checkpoints_dir_utils import get_latest_run_idrun_id = get_latest_run_id(experiment_name="<experiment_name>")
Combine with the above utility to fetch the path of the latest run directory.Next Steps:
- Dive into the checkpoints tutorial to grasp the essence of checkpoints, enabling you to resume trainings or access checkpoints from prior runs.
- The logs tutorial focuses on the log files stored in your run directories, offering insights into the training progression.
---
Documentation/Source/Experiment Monitoring (documentation/source/experiment_monitoring.md)
Third-party Experiment Monitoring
SuperGradients supports out-of-the-box Weights & Biases (wandb) and ClearML.
You can also inherit from our base class to integrate any monitoring tool with minimal code change.
Tensorboard
requirements: NoneTensorboard is natively integrated into the training and validation steps. You can find how to use it in this section.
DagsHub
[](https://colab.research.google.com/drive/11fW56pMpwOMHQSbQW6xxMRYvw1mEC-t-?usp=sharing)
requirements:
- Install
dagshub and mlflow
- You can set up DagsHub according to the official documentation, or you'll be guided interactively to sign in when you run the code with the logger
- Adapt your code like in the following examplepythonfrom super_gradients import Trainertrainer = Trainer("experiment_name")
model = ...
training_params = {
... # Your training params
"sg_logger": "dagshub_sg_logger", # DagsHub Logger, see class super_gradients.common.sg_loggers.dagshub_sg_logger.DagsHubSGLogger for details
"sg_logger_params": # Params that will be passes to __init__ of the logger super_gradients.common.sg_loggers.dagshub_sg_logger.DagsHubSGLogger
{
"dagshub_repository": "<REPO_OWNER>/<REPO_NAME>", # Optional: Your DagsHub project name, consisting of the owner name, followed by '/', and the repo name. If this is left empty, you'll be prompted in your run to fill it in manually.
"log_mlflow_only": False, # Optional: Change to true to bypass logging to DVC, and log all artifacts only to MLflow
"save_checkpoints_remote": True,
"save_tensorboard_remote": True,
"save_logs_remote": True,
}
}
trainer.train(model=model, training_params=training_params, ...)
Weights & Biases
requirements:- Install
wandb
- Set up wandb according to the official documentation
- Make sure to login (You can check if you have a ~/.netrc token)
- Adapt your code like in the following examplepythonfrom super_gradients import Trainertrainer = Trainer("experiment_name")
model = ...
training_params = {
... # Your training params
"sg_logger": "wandb_sg_logger", # Weights&Biases Logger, see class super_gradients.common.sg_loggers.wandb_sg_logger.WandBSGLogger for details
"sg_logger_params": # Params that will be passes to __init__ of the logger super_gradients.common.sg_loggers.wandb_sg_logger.WandBSGLogger
{
"project_name": "project_name", # W&B project name
"save_checkpoints_remote": True,
"save_tensorboard_remote": True,
"save_logs_remote": True,
"entity": "<YOUR-ENTITY-NAME>", # username or team name where you're sending runs
"api_server": "<OPTIONAL-WANDB-URL>" # Optional: In case your experiment tracking is not hosted at wandb servers
}
}
trainer.train(model=model, training_params=training_params, ...)
ClearML
requirements- Install
clearml
- Set up CleaML according to the official documentation
- Adapt your code like in the following examplepythonfrom super_gradients import Trainertrainer = Trainer("experiment_name")
model = ...
training_params = {
... # Your training params
"sg_logger": "clearml_sg_logger", # ClearML Logger, see class super_gradients.common.sg_loggers.wandb_sg_logger.ClearMLSGLogger for details
"sg_logger_params": # Params that will be passes to __init__ of the logger super_gradients.common.sg_loggers.wandb_sg_logger.ClearMLSGLogger
{
"project_name": "project_name", # ClearML project name
"save_checkpoints_remote": True,
"save_tensorboard_remote": True,
"save_logs_remote": True,
}
}
trainer.train(model=model, training_params=training_params, ...)
Integrate any other Monitoring tool
If your favorite monitoring tool is not supported by SuperGradients, you can simply implement a class inheriting from BaseSGLogger
that you will then pass to the training parameters.text/ Detailed source-code truncated for AI context efficiency. /
You can overwrite any method from BaseSGLogger to customize it to your need.
Then, you can pass it to your training_params exactly like WandB and ClearML.pythonfrom super_gradients import Trainertrainer = Trainer("experiment_name")
model = ...
training_params = {
..., # Your training params
"sg_logger": "CustomSGLogger", # Your custom CustomSGLogger
"sg_logger_params": {"project_name": "my_project_name"} # Params that will be passed to __init__ of your CustomSGLogger
}
trainer.train(model=model, training_params=training_params, ...)
Notes
-
@multi_process_safe prevents multiple training nodes to do the same action. Check out DDP documentation for more details
- @register_logger() registers your class into our factory, allowing it to be instantiated from a string.
- sg_logger_params only requires project_name, the rest is provided by the Trainer.
Uploading custom objects with a callback
Callbacks are the way to go when it comes to inserting small pieces of code into the training/validation loop of SuperGradients.
For more information, please check out our tutorial on how to use callbacks in SuperGradientsHere is a short example of how sg_logger can be used in callbacks:
pythonfrom super_gradients.training.utils.callbacks.base_callbacks import PhaseContext, Callbackdef do_something(inputs, target, preds):
pass
class DetectionVisualizationCallback2(Callback):
"""Save a custom metric to tensorboard and wandb/clearml"""
def __init__(self):
super(Callback, self).__init__()
def on_validation_batch_end(self, context: PhaseContext) -> None:
# Do something using the PhaseContext
custom_metric = do_something(context.inputs, context.target, context.preds)
# Save it to the tensorboard and wandb/clearml
context.sg_logger.add_scalar(
tag="custom_metric",
scalar_value=custom_metric,
global_step=context.epoch,
)
The sg_logger can also be used to upload files, text, images, checkpoints, ...
We encourage you to check out the API documentation of
super_gradients.common.sg_loggers.base_sg_logger.BaseSGLogger to see every available method.
Chose your monitoring tool in the recipes
You can update a recipe to use the monitoring tool you want by setting the sg_logger and sg_logger_params in recipes/training_hyperparams.Here is an example for WandB;
yamlsg_logger: wandb_sg_logger, # Weights&Biases Logger, see class super_gradients.common.sg_loggers.wandb_sg_logger.WandBSGLogger for details
sg_logger_params: # Params that will be passes to __init__ of the logger super_gradients.common.sg_loggers.wandb_sg_logger.WandBSGLogger
project_name: project_name, # W&B project name
save_checkpoints_remote: True,
save_tensorboard_remote: True,
save_logs_remote: True,
entity: <YOUR-ENTITY-NAME>, # username or team name where you're sending runs
api_server: <OPTIONAL-WANDB-URL> # Optional: In case your experiment tracking is not hosted at wandb servers
---
Documentation/Source/ImprovingTrainingTime (documentation/source/ImprovingTrainingTime.md)
Improving Training Time
Mixed Precision Training
Automatic mixed precision (AMP) is a feature in PyTorch that enables the use of lower-precision data types, such as float16, in deep learning models for improved memory and computation efficiency.
It automatically casts the model's parameters and buffers to a lower-precision data type, and dynamically rescales the activations to prevent underflow or overflow.
Most modern GPUs support float16 operations natively, and can therefore accelerate the training process.
To use
AMP in SuperGradients, you simply need to set mixed_precision=True in your training_hyperparams.In python script
pythonfrom super_gradients import Trainertrainer = Trainer("experiment_name")
model = ...
training_hyperparams = {"mixed_precision": True, ...:...}
trainer.train(model=model, training_hyperparams=training_hyperparams, ...)
In recipe
yamlmy_training_hyperparams.yaml
mixed_precision: True # Whether to use mixed precision or not.
Torch Compile
torch.compile API which can be used to improve the training time of the model.
This API can be used to fuse the operations in the model graph and optimize the model for the target device.SuperGradients support the
torch.compile API and can be used to improve the training time of the model.
Here we report the relative improvement (reduction) of training time for several models and tasks. We measure the training time of one epoch.
This includes iteration over training and validation datasets, loss & metric computation. Essentially all steps that are performed during training.
Please note that the improvement vary depending on the model architecture, dataset, and training hyperparameters.| Task | Recipe | Baseline (1 GPU) | Baseline (8 GPU) | 1 GPU With Compile | 8 GPU With Compile | Improvement, % (1 GPU) | Improvement, % (8 GPU) |
|-----------------------|---------------------------------|------------------|------------------|--------------------|--------------------|------------------------|------------------------|
| Semantic Segmentation | cityscapes_pplite_seg75 | 270.63 | 49.95 | 119.11 | 35.91 | 56% | 18% |
| Semantic Segmentation | cityscapes_regseg48 | 125.14 | 44.959 | 108.57 | 44.55 | 13.2% | 0.9% |
| Semantic Segmentation | cityscapes_segformer | 199.97 | 46.21 | 162.52 | 43.71 | 18.7% | 5.4% |
| Semantic Segmentation | cityscapes_stdc_seg75 | 425.19 | 73.07 | 153.16 | 45.89 | 63.9% | 37.19% |
| Semantic Segmentation | cityscapes_ddrnet | 226.51 | 51.78 | 174.11 | 48.29 | 23.1% | 7.3% |
| | | | | | | | |
| Object Detection | coco2017_yolo_nas_s | 1509 | 384.41 | 1379 | 376.10 | 8.6% | 2.42% |
| Object Detection | coco2017_yolo_nas_m | 2363 | 537.24 | 2090 | 508.40 | 11.5% | 0.19% |
| Object Detection | coco2017_yolo_nas_l | 3193 | 764.17 | 2869 | 745.58 | 10.14% | 2.43% |
| Object Detection | coco2017_ppyoloe_s/m/l/x | N/A | N/A | | N/A | N/A | N/A |
| Object Detection | coco2017_yolox_n/t/s/m/l/x | N/A | N/A | | N/A | N/A | N/A |
| Object Detection | coco2017_ssd_lite_mobilenet_v2 | N/A | N/A | | N/A | N/A | N/A |
| | | | | | | | |
| Classification | imagenet_efficientnet | | 425.61 | | 408.39 | | 4.1% |
| Classification | imagenet_mobilenetv3_large | | 373.73 | | 374.51 | | -0.2% |
| Classification | imagenet_regnetY | | 406.86 | | 383.04 | | 5.8% |
| Classification | imagenet_repvgg | | 407.19 | | 387.00 | | 4.9% |
| Classification | imagenet_resnet50 | | 481.36 | | 480.29 | | 0.22% |
| Classification | imagenet_vit_base | N/A | N/A | N/A | N/A | N/A | N/A |
| Classification | imagenet_vit_large | N/A | N/A | N/A | N/A | N/A | N/A |
In the table above, number are reported as speedup compared to the baseline training time.
Both experiments were run on 8x 3090 GPUs using PyTorch 2.0 with CUDA 11.8.
Training was done for 5 epochs and median value was picked to compute the speedup.
All experiments conducted with mixed precision (AMP) enabled, and SyncBN and EMA disabled.
Improvement percentage computed as follows:
100 * (baseline_time - compile_time) / baseline_time.To leverage use of compiled models in SuperGradients one need to pass the
torch_compile: True option to training hyperparameters:bashpython -m super_gradients.train_from_recipe --config-name=... training_hyperparams.torch_compile=True
In the YAML recipe:
yamlmy_training_recipe.yaml
training_hyperparams:
torch_compile: True
torch_compile_mode: default | reduce-overhead | max-autotune
Or programmatically:
pythonfrom super_gradients.training import Trainertrainer = Trainer(
...,
training_hyperparams = {
"torch_compile": True,
...
}
)
Avoiding common pitfalls:
* Don't use EMA with
torch.compile.
* Don't use SyncBN with torch.compile.
* You may need to reduce batch size during training by quite a lot (Up to 2x)
* Training with mixed precision gives the best performance boost.#### Exponential moving average
EMA and Torch CompileTorch Compile is still in its early stages and has some limitations. Not every model can be compiled.
Additionally, some training features can conflict with
torch.compile.
Here is what we found so far:* Exponential moving average
EMA is incompatible with torch.compile (At the moment of writing, this is true for SG release 3.1.2).
If you want to use torch.compile in your training, you need to disable EMA when using torch.compile.
yamltraining_hyperparams:
torch_compile: True
ema: False
#### Sync BatchNorm and Torch Compile In short:
When training using DDP, SyncBatchNorm layers _can be used_ with
torch.compile simultaneously.
Unfortunately, due to implementation details the Sync BN have to break the model graph each time BN layer is encountered to perform BN sync operation.
That means you will most likely get no speedup from torch.compile when using Sync BN.
It also was observed that it may require more GPU memory compared to training without torch.compile.
If you are running into errors when using torch.compile and SyncBatchNorm simultaneously, you can try lowering the batch size to see if that helps.#### Increased GPU memory consumption and Torch Compile
If during
torch.compile you are getting wierd CUDA-related exception messages you can try reducing batch size.
When using reduce-overhead or max-autotune modes peak GPU memory consumption may be higher compared to training without torch.compile.
It is good idea to run code with CUDA_LAUNCH_BLOCKING=1 first as it usually provides more meaningful error messages.#### Auto Mixed Precision and Torch Compile
Best speedup was achieved with combination of
torch.compile and AMP enabled. For F32 training torch.compile may not provide any speedup at all.
This is highly dependent on the target GPU and support of fp32 tensor cores, so we leave it up to the user to decide whether to use AMP or not.For AMP training with
torch.compile enabled, you need to pass mixed_precision: True to training hyperparameters:
yamltraining_hyperparams:
torch_compile: True
mixed_precision: True
---
Documentation/Source/Index (documentation/source/index.rst)
.. SuperGradients documentation master file, created by
sphinx-quickstart on Wed Nov 24 10:50:24 2021.
You can adapt this file completely to your liking, but it should at least
contain the root
toctree directive.Welcome to SuperGradients's documentation!
==========================================
.. toctree::
:maxdepth: 4
:caption: Welcome To SuperGradients
welcome
.. toctree::
:maxdepth: 4
:caption: Technical Documentation
super_gradients.common
super_gradients.training
.. toctree::
.. :maxdepth: 4
.. :caption: User Guide
Indices and tables
==================
* :ref:
genindex
* :ref:modindex
* :ref:search---
Documentation/Source/Installation (documentation/source/installation.md)
Installing SuperGradients
Requirements
General requirements
- Python 3.7, 3.8 or 3.9 installed.
- torch>=1.9.0
- https://pytorch.org/get-started/locally/
- The python packages that are specified in requirements.txt;
To train on nvidia GPUs
- Nvidia CUDA Toolkit >= 11.2
- CuDNN >= 8.1.x
- Nvidia Driver with CUDA >= 11.2 support (β₯460.x)Quick Installation
Install stable version using PyPi
See in PyPi
bashpip install super-gradients
That's it !
Important: If PyTorch was not already installed on your environment, you might need to reinstall
a Pytorch version suitable for your CUDA version. Go into PyTorch installation page
and follow the instructions to install the correct version.
Install using GitHub</summary>
bashpip install git+https://github.com/Deci-AI/super-gradients.git@stable
---
Documentation/Source/KD (documentation/source/KD.md)
Knowledge Distillation (KD)
Pre-requisites: Training in SG, Training with Configuration Files
Knowledge distillation is a technique in deep learning that aims to transfer the knowledge of a large, pre-trained neural network model (the "teacher") to a smaller, more computationally efficient model (the "student"). This is accomplished by training the student to mimic the teacher's predictions and the ground-truth labels. The student network can also be designed to have a different architecture from the teacher, making it possible to distill the knowledge of a complex teacher network into a lighter and faster student network for deployment in real-world applications.
The training flow with Knowledge distillation in SG is similar to regular training. For standard training, we used SGs
Trainer class - which was in charge of training the model, evaluating test data, making predictions, and saving checkpoints.
Equivalently, for knowledge distillation, we use the KDTrainer class which inherits from Trainer.If for regular training with
Trainer, the general flow is:python...
trainer = Trainer("my_experiment")
train_dataloader = ...
valid_dataloader = ...
model = ...train_params = {...}
trainer.train(model=model, training_params=train_params, train_loader=train_dataloader, valid_loader=valid_dataloader)
Then for training with knowledge distillation, the general flow is:
pythonfrom super_gradients.training.kd_trainer import KDTrainer
...
kd_trainer = KDTrainer("my_experiment")
train_dataloader = ...
valid_dataloader = ...student_model = ...
teacher_model = ...
train_params = {...}
kd_trainer.train(student=student_model, teacher=teacher_model, training_params=train_params, train_loader=train_dataloader, valid_loader=valid_dataloader)
Check out our knowledge distillation tutorial notebook to see a practical example.
Knowledge Distillation Training: Key Components
KDModule
The most apparent difference in the training flow using knowledge distillation is that it requires two networks: the "teacher" and the "student".
The relation between the two is also configurable - for example, we may decide that the teacher model should preprocess the inputs differently.
For that matter, SG introduces a new
torch.nn.Module that wraps both the student and the teacher models: KDModule.
Upon calling KDTrainer.train(), the teacher and student models are passed along the kd_arch_params to initialize a KDModule instance.Passing a
KDModule instance explicitly to KDTrainer.train() through the model argument instead of student and teacher models is also possible, which gives the users the option to customize KD to their needs.A high-level example of KD customization:
pythonimport torch.nn
from super_gradients.training.kd_trainer import KDTrainer...
class MyKDModule(KDModule):
...
def forward(self, x: torch.Tensor)->KDOutput:
intermediate_output_student = self.student.extract_intermediate_output(x, layer_ids=[1, 3, -1])
intermediate_output_teacher = self.teacher.extract_intermediate(x, layer_ids=[1, 3, -1])
return KDOutput(student_output=intermediate_output_student, teacher_output=intermediate_output_teacher)
class MyKDLoss(torch.nn.Module):
...
def forward(self, preds: KDOutput, target: torch.Tensor):
# does something with the intermediate outputs
...
kd_trainer = KDTrainer("my_customized_kd_experiment")
train_dataloader = ...
valid_dataloader = ...
student_model = ...
teacher_model = ...
kd_model = MyKDModule(student=student_model, teacher=teacher_model)
train_params = {'loss': MyKDLoss(),
...}
kd_trainer.train(model=kd_model, training_params=train_params,
train_loader=train_dataloader, valid_loader=valid_dataloader)
KDOutput
KDOutput defines the structure of the output of KDModule and has two self-explanatory attributes: student_output and teacher_output.
KDTrainer uses these attributes behind the scenes to perform the usual operations of regular training, such as metrics calculations.
This means that when customizing KD, it's essential for the custom KDModule to stick to this output format.KD Losses
Currently, KDLogitsLoss is currently the only supported loss function in SGs KD losses bank, but more is to come.
Note that during KD training, the
KDModule outputs (which are of KDOutput instance) are passed to the loss's forward method as predictions.Knowledge Distillation Training: Checkpoints
Checkpointing during KD training is generally the same as checkpointing without KD.
Nevertheless, there are a few differences worth mentioning:
-
ckpt_latest.pth contains the state dict of the entire KDModule.
- ckpt_best.pth contains the state dict of the student only.
- When training with EMA, ckpt_best.pths net entry holds the EMA network.
Knowledge Distillation Training with Configuration Files
As done when training without knowledge distillation, to train with configuration files, we call the
KDTrainer.train_from_config method, which assumes a specific configuration structure.
When training with KD, the same structure and required fields hold, but we introduce a few additions:-
arch_params are being passed to the KDModule constructor. For example, in our Resnet50 KD training on Imagenet, we handle the difference in preprocessing of the teacher, which expects different normalization by passing the KDModule a normalization adaptor module:
yamlsuper_gradients/recipes/imagenet_resnet50_kd.yaml
...
arch_params:
teacher_input_adapter:
_target_: super_gradients.training.utils.kd_trainer_utils.NormalizationAdapter
mean_original: [0.485, 0.456, 0.406]
std_original: [0.229, 0.224, 0.225]
mean_required: [0.5, 0.5, 0.5]
std_required: [0.5, 0.5, 0.5]
Warning: Remember to distinguish the arch params being passed to the KDModule constructor from the student ones.
-
student_architecture, teacher_architecture, student_arch_params, student_checkpoint_params, teacher_arch_params, and teacher_checkpoint_params play the same role as architecture, arch_params and checkpoint_params for instantiating our model in non-KD training, and are being passed to models.get(...) to instantiate the teacher and the student:yaml...
student_architecture: resnet50
teacher_architecture: beit_base_patch16_224student_arch_params:
num_classes: 1000
teacher_arch_params:
num_classes: 1000
image_size: [224, 224]
patch_size: [16, 16]
teacher_checkpoint_params:
...
pretrained_weights: imagenet
student_checkpoint_params:
...
Any KD recipe can be launched with our train_from_kd_recipe_example script.
---
Documentation/Source/LICENSE (documentation/source/LICENSE.md)
Apache License
Version 2.0, January 2004
http://www.apache.org/licenses/
TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
1. Definitions.
"License" shall mean the terms and conditions for use, reproduction,
and distribution as defined by Sections 1 through 9 of this document.
"Licensor" shall mean the copyright owner or entity authorized by
the copyright owner that is granting the License.
"Legal Entity" shall mean the union of the acting entity and all
other entities that control, are controlled by, or are under common
control with that entity. For the purposes of this definition,
"control" means (i) the power, direct or indirect, to cause the
direction or management of such entity, whether by contract or
otherwise, or (ii) ownership of fifty percent (50%) or more of the
outstanding shares, or (iii) beneficial ownership of such entity.
"You" (or "Your") shall mean an individual or Legal Entity
exercising permissions granted by this License.
"Source" form shall mean the preferred form for making modifications,
including but not limited to software source code, documentation
source, and configuration files.
"Object" form shall mean any form resulting from mechanical
transformation or translation of a Source form, including but
not limited to compiled object code, generated documentation,
and conversions to other media types.
"Work" shall mean the work of authorship, whether in Source or
Object form, made available under the License, as indicated by a
copyright notice that is included in or attached to the work
(an example is provided in the Appendix below).
"Derivative Works" shall mean any work, whether in Source or Object
form, that is based on (or derived from) the Work and for which the
editorial revisions, annotations, elaborations, or other modifications
represent, as a whole, an original work of authorship. For the purposes
of this License, Derivative Works shall not include works that remain
separable from, or merely link (or bind by name) to the interfaces of,
the Work and Derivative Works thereof.
"Contribution" shall mean any work of authorship, including
the original version of the Work and any modifications or additions
to that Work or Derivative Works thereof, that is intentionally
submitted to Licensor for inclusion in the Work by the copyright owner
or by an individual or Legal Entity authorized to submit on behalf of
the copyright owner. For the purposes of this definition, "submitted"
means any form of electronic, verbal, or written communication sent
to the Licensor or its representatives, including but not limited to
communication on electronic mailing lists, source code control systems,
and issue tracking systems that are managed by, or on behalf of, the
Licensor for the purpose of discussing and improving the Work, but
excluding communication that is conspicuously marked or otherwise
designated in writing by the copyright owner as "Not a Contribution."
"Contributor" shall mean Licensor and any individual or Legal Entity
on behalf of whom a Contribution has been received by Licensor and
subsequently incorporated within the Work.
2. Grant of Copyright License. Subject to the terms and conditions of
this License, each Contributor hereby grants to You a perpetual,
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
copyright license to reproduce, prepare Derivative Works of,
publicly display, publicly perform, sublicense, and distribute the
Work and such Derivative Works in Source or Object form.
3. Grant of Patent License. Subject to the terms and conditions of
this License, each Contributor hereby grants to You a perpetual,
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
(except as stated in this section) patent license to make, have made,
use, offer to sell, sell, import, and otherwise transfer the Work,
where such license applies only to those patent claims licensable
by such Contributor that are necessarily infringed by their
Contribution(s) alone or by combination of their Contribution(s)
with the Work to which such Contribution(s) was submitted. If You
institute patent litigation against any entity (including a
cross-claim or counterclaim in a lawsuit) alleging that the Work
or a Contribution incorporated within the Work constitutes direct
or contributory patent infringement, then any patent licenses
granted to You under this License for that Work shall terminate
as of the date such litigation is filed.
4. Redistribution. You may reproduce and distribute copies of the
Work or Derivative Works thereof in any medium, with or without
modifications, and in Source or Object form, provided that You
meet the following conditions:
(a) You must give any other recipients of the Work or
Derivative Works a copy of this License; and
(b) You must cause any modified files to carry prominent notices
stating that You changed the files; and
(c) You must retain, in the Source form of any Derivative Works
that You distribute, all copyright, patent, trademark, and
attribution notices from the Source form of the Work,
excluding those notices that do not pertain to any part of
the Derivative Works; and
(d) If the Work includes a "NOTICE" text file as part of its
distribution, then any Derivative Works that You distribute must
include a readable copy of the attribution notices contained
within such NOTICE file, excluding those notices that do not
pertain to any part of the Derivative Works, in at least one
of the following places: within a NOTICE text file distributed
as part of the Derivative Works; within the Source form or
documentation, if provided along with the Derivative Works; or,
within a display generated by the Derivative Works, if and
wherever such third-party notices normally appear. The contents
of the NOTICE file are for informational purposes only and
do not modify the License. You may add Your own attribution
notices within Derivative Works that You distribute, alongside
or as an addendum to the NOTICE text from the Work, provided
that such additional attribution notices cannot be construed
as modifying the License.
You may add Your own copyright statement to Your modifications and
may provide additional or different license terms and conditions
for use, reproduction, or distribution of Your modifications, or
for any such Derivative Works as a whole, provided Your use,
reproduction, and distribution of the Work otherwise complies with
the conditions stated in this License.
5. Submission of Contributions. Unless You explicitly state otherwise,
any Contribution intentionally submitted for inclusion in the Work
by You to the Licensor shall be under the terms and conditions of
this License, without any additional terms or conditions.
Notwithstanding the above, nothing herein shall supersede or modify
the terms of any separate license agreement you may have executed
with Licensor regarding such Contributions.
6. Trademarks. This License does not grant permission to use the trade
names, trademarks, service marks, or product names of the Licensor,
except as required for reasonable and customary use in describing the
origin of the Work and reproducing the content of the NOTICE file.
7. Disclaimer of Warranty. Unless required by applicable law or
agreed to in writing, Licensor provides the Work (and each
Contributor provides its Contributions) on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
implied, including, without limitation, any warranties or conditions
of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
PARTICULAR PURPOSE. You are solely responsible for determining the
appropriateness of using or redistributing the Work and assume any
risks associated with Your exercise of permissions under this License.
8. Limitation of Liability. In no event and under no legal theory,
whether in tort (including negligence), contract, or otherwise,
unless required by applicable law (such as deliberate and grossly
negligent acts) or agreed to in writing, shall any Contributor be
liable to You for damages, including any direct, indirect, special,
incidental, or consequential damages of any character arising as a
result of this License or out of the use or inability to use the
Work (including but not limited to damages for loss of goodwill,
work stoppage, computer failure or malfunction, or any and all
other commercial damages or losses), even if such Contributor
has been advised of the possibility of such damages.
9. Accepting Warranty or Additional Liability. While redistributing
the Work or Derivative Works thereof, You may choose to offer,
and charge a fee for, acceptance of support, warranty, indemnity,
or other liability obligations and/or rights consistent with this
License. However, in accepting such obligations, You may act only
on Your own behalf and on Your sole responsibility, not on behalf
of any other Contributor, and only if You agree to indemnify,
defend, and hold each Contributor harmless for any liability
incurred by, or claims asserted against, such Contributor by reason
of your accepting any such warranty or additional liability.
END OF TERMS AND CONDITIONS
APPENDIX: How to apply the Apache License to your work.
To apply the Apache License to your work, attach the following
boilerplate notice, with the fields enclosed by brackets "[]"
replaced with your own identifying information. (Don't include
the brackets!) The text should be enclosed in the appropriate
comment syntax for the file format. We also recommend that a
file or class name and description of purpose be included on the
same "printed page" as the copyright notice for easier
identification within third-party archives.
Copyright [2022] [Deci-AI]
Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.
---
Documentation/Source/Logs (documentation/source/logs.md)
Local Logging
SuperGradients automatically logs multiple files locally that can help you explore your experiments results.
This includes 1 tensorboard and 3 .txt files.
Absolutely. I understand your requirements. Here's a more concise and structured introduction:
Directory Structure Overview:
- ckpt_root_dir: The root directory where all experiments are stored.
- experiment_name: The specific folder dedicated to your current experiment.
- run_dir: Unique identifier for each training run; contains all associated checkpoints and logs.For a deeper dive into checkpoints, visit our detailed guide.
I. Tensorboard logging
To easily keep track of your experiments, SuperGradients saves your results in events.out.tfevents format that can be used by tensorboard.What does it include? This tensorboard includes all of your training and validation metrics but also other information such as learning rate, system metrics (CPU, GPU, ...), and more.
Where is it saved?
<ckpt_root_dir>/<experiment_name>/<run_dir>/events.out.tfevents.<unique_id>How to launch?
tensorboard --logdir <ckpt_root_dir>/<experiment_name>/<run_dir>II. Experiment logging
In case you cannot launch a tensorboard instance, you can still find a summary of your experiment saved in a readable .txt format.What does it include? The experiment configuration and training/validation metrics.
Where is it saved?
<ckpt_root_dir>/<experiment_name>/<run_dir>/experiment_logs_<date>.txtIII. Console logging
For better debugging and understanding of past runs, SuperGradients gathers all the print statements and logs into a
local file, providing you the convenience to review console outputs of any experiment at any time.What does it include? All the prints and logs that were displayed on the console, but not the filtered logs.
Where is it saved?
- Upon importing SuperGradients, console outputs and logs will be stored in
~/sg_logs/console.log.
- When instantiating the super_gradients.Trainer, all console outputs and logs will be redirected to the experiment folder <ckpt_root_dir>/<experiment_name>/<run_dir>/console_<date>.txt.How to set log level? You can filter the logs displayed on the console by setting the environment variable
CONSOLE_LOG_LEVEL=<LOG-LEVEL> # DEBUG/INFO/WARNING/ERRORIV. Loggers logging
Contrary to the console logging, the logger logging is restricted to the loggers messages (such as logger.log, logger.info, ...).
This means that it includes any log that was under the logging level (logging.DEBUG for instance), but not the prints.What does it include? Anything logged with a logger (
logger.log, logger.info, ...), even the filtered logs.Where is it saved?
<ckpt_root_dir>/<experiment_name>/<run_dir>/logs_<date>.txtHow to set log level? You can filter the logs saved in the file by setting the environment variable
FILE_LOG_LEVEL=<LOG-LEVEL> # DEBUG/INFO/WARNING/ERROR
(Additional) Hydra config folder
Only when training using hydra recipe.What does it include?
text<ckpt_root_dir>/<experiment_name>/<run_dir>/
ββ .hydra
ββconfig.yaml # A single config file that regroups the config files used to run the experiment
ββhydra.yaml # Some Hydra metadata
ββoverrides.yaml # Any override passed after --config-name=<config-name>
SUMMARY
text<ckpt_root_dir>/<experiment_name>/<run_dir>/
ββ ... (all the model checkpoints)
ββ events.out.tfevents.<unique_id> # Tensorboard artifact
ββ experiment_logs_<date>.txt # Config and metrics related to experiment
ββ console_<date>.txt # Logs and prints that were displayed in the users console
ββ logs_<date>.txt # Every log
ββ .hydra # (Additional) If experiment launched from a recipe:
ββconfig.yaml # A single config file that regroups the config files used to run the experiment
ββhydra.yaml # Some Hydra metadata
ββoverrides.yaml # Any override passed after --config-name=<config-name>
Other
#### Environment Sanity Check
SuperGradients automatically checks compatibility between the installed libraries and the required ones.
It will log an error - but not stop the code - for each library that was installed with a version lower than required.
For libraries with version higher than required, this information will just be logged at a DEBUG level.
#### Crash Tip
It can sometimes be very time consuming to debug an exceptions when the error raised is not explicit.
To avoid this, SuperGradients implemented a Crash Tip system that decorates errors raised from different libraries to help you fix the issue.
Example
The error raised by hydra when you made an indentation error is hard to understand (see topmost RuntimeError).
Under the exception, SuperGradients prints a Crash Tip that explains what went wrong, and how to fix it.
The number of crash tips is limited to cases that were faced by the community, so if you face an exception that is hard to understand feel free to share with us!
How to disable? The Crash tip can be shut down by setting the environment variable
CRASH_HANDLER=FALSE.---
Documentation/Source/Losses (documentation/source/Losses.md)
Losses
SuperGradients can support any PyTorch-based loss function. Additionally, multiple Loss function implementations for various tasks are also supported:
CrossEntropyLoss
MSE
RSquaredLoss
ShelfNetOHEMLoss
ShelfNetSemanticEncodingLoss
YoloXDetectionLoss
YoloXFastDetectionLoss
SSDLoss
STDCLoss
BCEDiceLoss
KDLogitsLoss
DiceCEEdgeLoss
All the above, are just string aliases for the underlying torch.nn.Module classes, implementing the specified loss functions.
Basic Usage of Implemented Loss Functions
The most basic use case is when using a direct Trainer.train(...) call:
In your
my_training_script.py:
python...
trainer = Trainer("external_criterion_test")
train_dataloader = ...
valid_dataloader = ...
model = ...train_params = {
...
"loss": "CrossEntropyLoss",
"criterion_params": {}
...
}
trainer.train(model=model, training_params=train_params, train_loader=train_dataloader, valid_loader=valid_dataloader)
Note that object names in SG are not case-sensitive nor symbol-sensitive, so "CrossEntropy could have been passed as well.
Since most IDEs support auto-completion, for your convenience, you can use our object_names module:
pythonfrom super_gradients.common.object_names import Losses
Then simply instead of "CrossEntropyLoss", use
pythonLosses.CROSS_ENTROPY
Another use case is when using configuration files. For example, when training using train_from_recipe (or similar, when the underlying train method that is being called is Trainer.train_from_config(...)).
When doing so, in your
my_training_hyperparams.yaml file:
yaml...loss: YoloXDetectionLoss
criterion_params:
strides: [8, 16, 32] # output strides of all yolo outputs
num_classes: 80
Note that two
training_params parameters define the loss function: loss which defines the type of the loss, andcriterion_params dictionary which will be unpacked to the underlying YoloXDetectionLoss class constructor.Passing Instantiated nn.Module Objects as Loss Functions
SuperGradients also supports passing instantiated nn.Module Objects as demonstrated below:
When using a direct Trainer.train(...) call, in your
my_training_script.py simply pass the instantiated nn.Module under the "loss" key inside training_params:
python...
trainer = Trainer("external_criterion_test")
train_dataloader = ...
valid_dataloader = ...
model = ...train_params = {
...
"loss": torch.nn.CrossEntropy()
...
}
trainer.train(model=model, training_params=train_params, train_loader=dataloader, valid_loader=dataloader)
Though not as convenient as using
register_loss (discussed further into detail in the next sub-section), one can also equivalently instantiate objects when using train_from_recipe (or similar, when the underlying train method is Trainer.train_from_config(...) as demonstrated below:
In your
my_training_hyperparams.yaml file:
yaml...loss:
_target_: torch.nn.CrossEntropy
Note that when passing an instantiated loss object, criterion_params will be ignored.
Using Your Own Loss
SuperGradients also supports user-defined loss functions assuming they are torch.nn.Module inheritors, and that their
forward signature is in the form:pythonimport torch.nnMyLoss(torch.nn.Module):
...
forward(preds, target):
...
And as the argument names suggest, the first argument is the model's output, and target is the label/ground truth (argument naming is arbitrary and does not need to be specifically 'preds' or 'target').
Loss functions accepting additional arguments in their forward method will be supported in the future.Using Your Own Loss - Logging Loss Outputs
In the most common case, where the loss function returns a single item for backprop the loss output will appear in
the logs, training logs (i.e Tensorboards and any other supported SGLogger, for more information on SGLoggers click here), over epochs under <LOSS_CLASS.__name__>.
forward(...) should return a (loss, loss_items) tuple where loss is the tensor used
for backprop (i.e what your original loss function returns), and loss_items should be a tensor of
shape (n_items) consisting of values computed during the forward pass which we desire to log over the
entire epoch. For example- the loss itself should always be logged. Another example is a scenario
where the computed loss is the sum of a few components we would like to log.
For example:
pythonclass MyLoss(_Loss):
...
def forward(self, inputs, targets):
...
total_loss = comp1 + comp2
loss_items = torch.cat((total_loss.unsqueeze(0),comp1.unsqueeze(0), comp2.unsqueeze(0)).detach()
return total_loss, loss_items
train_params = {
...,
"loss": MyLoss(),
"metric_to_watch": "MyLoss/loss_0"
}
Trainer.train(
...,
train_params=train_params
)
The above snippet will log
MyLoss2/loss_0, MyLoss2/loss_1 and MyLoss2/loss_2 as they have been named by their positional index in loss_items.
Note we also defined "MyLoss2/loss_0" to be our watched metric which means we save our checkpoint every epoch we reach the best loss score.For more visibility, you can also set a "component_names" property in the loss class,
to be a list of strings, of length n_items whose ith element is the name of the ith entry in loss_items.
Then each item will be logged, rendered on the tensorboard, and "watched" (i.e saving model checkpoints
according to it) under
<LOSS_CLASS.__name__>/<COMPONENT_NAME>. For example:
pythonclass MyLoss(_Loss):
...
def forward(self, inputs, targets):
...
total_loss = comp1 + comp2
loss_items = torch.cat((total_loss.unsqueeze(0),comp1.unsqueeze(0), comp2.unsqueeze(0)).detach()
return total_loss, loss_items
...
@property
def component_names(self):
return ["total_loss", "my_1st_component", "my_2nd_component"]train_params = {
...,
"loss": MyLoss(),
"metric_to_watch": "MyLoss/my_1st_component"
}
Trainer.train(
...,
train_params=train_params
)
The above code will log and monitor
MyLoss/total_loss, MyLoss/my_1st_component and MyLoss/my_2nd_component.
Since running logs will save the loss_items in some internal state, it is recommended to
detach loss_items from their computational graph for memory efficiency.
Using Your Own Loss - Training with Configuration Files
When using configuration files, for example, training using train_from_recipe (or similar, when the underlying train method that is being called is Trainer.train_from_config(...)), In your
my_loss.py, register your loss class by decorating the class with register_loss:
pythonimport torch.nn
from super_gradients.common.registry import register_loss
@register_loss("my_loss")
class MyLoss(torch.nn.Module):
...
Then, in your
my_training_hyperparams.yaml, use "my_loss" in the same way as any other loss supported in SG:
yaml...loss: my_loss
criterion_params:
...
Last, in your
my_train_from_recipe_script.py file, just import the newly registered class (even though the class itself is unused, just to trigger the registry):
pythonfrom omegaconf import DictConfig
import hydra
import pkg_resources
from my_loss import MyLoss
from super_gradients import Trainer, init_trainer
@hydra.main(config_path=pkg_resources.resource_filename("super_gradients.recipes", ""), version_base="1.2")
def main(cfg: DictConfig) -> None:
Trainer.train_from_config(cfg)
def run():
init_trainer()
main()
if __name__ == "__main__":
run()
---
Documentation/Source/LRAssignment (documentation/source/LRAssignment.md)
Assigning Learning Rates in SG
The initial_lr training hyperparameter allows you to specify different learning rates for different layers or groups of parameters in your neural network. This can be particularly useful for fine-tuning pre-trained models or when different parts of your model require different learning rate settings for optimal training.Using
initial_lr as a Scalar:
When initial_lr is a single floating-point number, it sets a uniform learning rate for all model parameters. For example, initial_lr = 0.01:pythonDefine training parameters
training_params = {
"initial_lr": 0.01,
"loss": "cross_entropy",
# ... other training parameters
}Initialize the Trainer
trainer = Trainer("simple_net_training")Define model
model = Define data loaders
train_dataloader = ...
test_dataloader = ...Train the model
trainer.train(model, training_params, train_dataloader, test_dataloader)
Using
initial_lr as a Mapping:
initial_lr can also be a mapping where keys are the prefixes of the named parameters of the model, and values are the learning rates
for those specific groups. This approach offers granular control over the learning rates for different parts of the model.* Each key in the
initial_lr dictionary acts as a prefix to match the named parameters in the model. The learning rate associated with a key is applied to all parameters whose names start with that prefix.* The "default" key is essential, as it provides a fallback learning rate for any parameter that does not match other specified prefixes.
* Freezing parameters can be done by assigning a learning rate of 0 to a specific prefix. By doing so, you will be preventing them from being updated during training.
For example, in the below snippet
conv1 and conv2 will be frozen, and fc1 and fc2 will be trained with an initial learning rate of 0.001:pythonclass SimpleNet(nn.Module):
def __init__(self):
super(SimpleNet, self).__init__()
self.conv1 = nn.Conv2d(1, 20, 5)
self.conv2 = nn.Conv2d(20, 50, 5)
self.fc1 = nn.Linear(50 4 4, 500)
self.fc2 = nn.Linear(500, 10) def forward(self, x):
x = nn.functional.relu(self.conv1(x))
x = nn.functional.relu(self.conv2(x))
x = x.view(-1, 50 4 4)
x = nn.functional.relu(self.fc1(x))
x = self.fc2(x)
return x
trainer = Trainer("simple_net_training")
Define model
model = SimpleNet()Define data loaders
train_dataloader = ...
test_dataloader = ...Define training parameters
training_params = {
"initial_lr": {"conv": 0.001, "default": 0.},
"loss": "cross_entropy",
# ... other training parameters
}Train the model
trainer.train(model, training_params, train_dataloader, test_dataloader)
Fine-Tuning with the
finetune Feature The
finetune parameter in SG adds another layer of control for model training. When set to True, it enables selective freezing of parts of the model, a technique often used in fine-tuning pre-trained models.
This feature is supported for all models in the SG model zoo that implement the get_finetune_lr_dict method. It is useful when one is not familiar with the different parts of the network.For example, in the below the detection heads of YoloNAS will be trained with an initial learning rate of 0.01 while the rest of the network is frozen:
pythontrainer = Trainer("simple_net_training")Define model
model = models.get(Models.YOLO_NAS_S, pretrained_weights="coco", num_classes=2)Define data loaders
train_dataloader = ...
test_dataloader = ...Define training parameters
training_params = {
"initial_lr": 0.01,
"finetune": True
# ... other training parameters
}Train the model
trainer.train(model, training_params, train_dataloader, test_dataloader)
How
finetune Works - When
finetune is set to True, the model automatically freezes a part of itself based on the definitions in the get_finetune_lr_dict method.
- The get_finetune_lr_dict method returns a dictionary mapping learning rates to the unfrozen part of the network, in the same fashion as when initial_lr is used as a mapping.
For example, the implementation for YoloNAS:
pythondef get_finetune_lr_dict(self, lr: float):
return {"heads": lr, "default": 0}
- If initial_lr is already a mapping, using finetune will raise an error. It's designed to work when initial_lr is unset or a float.---
Documentation/Source/LRScheduling (documentation/source/LRScheduling.md)
Learning Rate Scheduling
When training deep neural networks, it is often useful to reduce learning rate as the training progresses. This can be done by using pre-defined learning rate schedules or adaptive learning rate methods.
Learning rate scheduling type is controlled by the training parameter
lr_mode. From Trainer.train(...) docs:
lr_mode : Union[str, Mapping] When str:
Learning rate scheduling policy, one of ['StepLRScheduler','PolyLRScheduler','CosineLRScheduler','FunctionLRScheduler'].
'StepLRScheduler' refers to constant updates at epoch numbers passed through
lr_updates. Each update decays the learning rate by lr_decay_factor. 'CosineLRScheduler' refers to the Cosine Anealing policy as mentioned in https://arxiv.org/abs/1608.03983. The final learning rate ratio is controlled by
cosine_final_lr_ratio training parameter. 'PolyLRScheduler' refers to the polynomial decrease: in each epoch iteration
self.lr = self.initial_lr * pow((1.0 - (current_iter / max_iter)), 0.9) 'FunctionLRScheduler' refers to a user-defined learning rate scheduling function, that is passed through
lr_schedule_function.For example, the training code below will start with an initial learning rate of 0.1 and decay by 0.1 at epochs 100,150 and 200:
pythonfrom super_gradients.training import Trainer
trainer = Trainer("my_custom_scheduler_training_experiment")
train_dataloader = ...
valid_dataloader = ...
model = ...
train_params = {
"initial_lr": 0.1,
"lr_mode":"StepLRScheduler",
"lr_updates": [100, 150, 200],
"lr_decay_factor": 0.1,
...,
}
trainer.train(model=model, training_params=train_params, train_loader=train_dataloader, valid_loader=valid_dataloader)
<details>
<summary>Equivalent in a <code>.yaml</code> configuration file:</summary>
yamltraining_hyperparams:
initial_lr: 0.1
lr_mode: StepLRScheduler
user_lr_updates:
- 100
- 150
- 200
lr_decay_factor: 0.1
......
</details>
Using Custom LR Schedulers
Prerequisites: phase callbacks, training with configuration files.
In SG, learning rate schedulers are implemented as phase callbacks.
They read the learning rate from the
PhaseContext in their __call__ method, calculate the new learning rate according to the current state of training, and update the optimizer's param groups.For example, the code snippet from the previous section translates "lr_mode":"StepLRScheduler" to a
super_gradients.training.utils.callbacks.callbacks.StepLRScheduler instance, which is added to the phase callbacks list.Implementing Your Own Scheduler
A custom learning rate scheduler should inherit from LRCallbackBase, so let's take a look at it:text/ Detailed source-code truncated for AI context efficiency. /
So when writing a custom scheduler, we need to override two methods:
1.
perform_scheduling: This is where the new learning rate is calculated. The lr attribute is updated according. Then, in order to update the optimizer's parameter groups a call for update_lr should also be done (or update the optimizers parameter groups with your own logic explicitly).
2. is_lr_scheduling_enabled: Predicate that controls whether to perform lr scheduling based on values in context.We will demonstrate how this is done by implementing a simple scheduler that decays the learning rate by a user-defined rate at user-defined epoch numbers.
pythonfrom super_gradients.training.utils.callbacks import LRCallbackBase, Phase
from super_gradients.common.abstractions.abstract_logger import get_loggerlogger = get_logger(__name__)
class UserStepLRCallback(LRCallbackBase):
def __init__(self, lr_updates: list, lr_decay_factors: list, kwargs):
super(UserStepLRCallback, self).__init__(Phase.TRAIN_EPOCH_END, kwargs)
assert len(lr_updates) == len(lr_decay_factors)
self.lr_updates = lr_updates
self.lr_decay_factors = lr_decay_factors
def perform_scheduling(self, context):
curr_lr = self.initial_lr
for epoch_idx, epoch_decay_rate in zip(self.lr_updates, self.lr_decay_factors):
if epoch_idx <= context.epoch:
curr_lr *= epoch_decay_rate
self.lr = curr_lr
self.update_lr(context.optimizer, context.epoch, None)
def is_lr_scheduling_enabled(self, context):
return self.training_params.lr_warmup_epochs <= context.epoch
Notes
- We specified that scheduling is enabled only after
lr_warmup_epochs, this means that during lr warmup no updates will be done, even if such epoch is specifed!
- Notice the Phase.TRAIN_EPOCH_END which we pass to the constructor, this means that our __call__ is triggered inside on_train_loader_end(self, context) (see new callbacks API mapping between Phase to Callback methods.)Now, we need to register our new scheduler so we can pass it through the
lr_mode training parameter.
First we decorate our class with the register_lr_scheduler.
pythonmyscheduler.py
from super_gradients.training.utils.callbacks import LRCallbackBase, Phase
from super_gradients.common.abstractions.abstract_logger import get_logger
from super_gradients.common.registry import register_lr_scheduler
logger = get_logger(__name__)
@register_lr_scheduler("user_step")
class UserStepLRCallback(LRCallbackBase):
def __init__(self, user_lr_updates: list, user_lr_decay_factors: list, kwargs):
super(UserStepLRCallback, self).__init__(Phase.TRAIN_EPOCH_END, kwargs)
assert len(user_lr_updates) == len(user_lr_decay_factors)
self.lr_updates = user_lr_updates
self.lr_decay_factors = user_lr_decay_factors
def perform_scheduling(self, context):
curr_lr = self.initial_lr
for epoch_idx, epoch_decay_rate in zip(self.lr_updates, self.lr_decay_factors):
if epoch_idx <= context.epoch:
curr_lr *= epoch_decay_rate
self.lr = curr_lr
self.update_lr(context.optimizer, context.epoch, None)
def is_lr_scheduling_enabled(self, context):
return self.training_params.lr_warmup_epochs <= context.epoch
Next, simply import it (even if the class itself isn't used on the training script code page) to trigger the registry.
pythonmy_train_script.py
from super_gradients.training import Trainer
from myscheduler import UserStepLRCallback # triggers registry, now we can pass "lr_mode": "user_step"
...
And finally, use your new scheduler just as any other one supported by SG.
pythontrainer = Trainer("my_custom_scheduler_training_experiment")The following code sections marked with '...' are placeholders
indicating additional necessary code that is not shown for simplicity.
train_dataloader = ...
valid_dataloader = ...
model = ...train_params = {
"initial_lr": 0.1,
"lr_mode": "user_step",
"user_lr_updates": [100, 150, 200], # WILL BE PASSED TO UserStepLRCallback CONSTRUCTOR
"user_lr_decay_factors": [0.1, 0.01, 0.001], # WILL BE PASSED TO UserStepLRCallback CONSTRUCTOR
...
}
trainer.train(model=model, training_params=train_params, train_loader=train_dataloader, valid_loader=valid_dataloader)
Note that internally, Trainer unpacks training_params to the scheduler callback constructor, so we pass scheduler related parameters through training_params as well.
<details>
<summary>Equivalent in a <code>.yaml</code> configuration file:</summary>
yamltraining_hyperparams:
initial_lr: 0.1
lr_mode: user_step
user_lr_updates: # WILL BE PASSED TO UserStepLRCallback CONSTRUCTOR
- 100
- 150
- 200
user_lr_decay_factors: # WILL BE PASSED TO UserStepLRCallback CONSTRUCTOR
- 0.1
- 0.01
- 0.001
......
</details>Using PyTorchs Native LR Schedulers (torch.optim.lr_scheduler)
PyTorch offers a wide variety of learning rate schedulers.
They can all be easily used by passing a Mapping through the lr_mode parameter, following aa simple API.
From
Trainer.train(...) docs: When Mapping, refers to a torch.optim.lr_scheduler._LRScheduler, following the below API:
lr_mode = {LR_SCHEDULER_CLASS_NAME: {LR_SCHEDULER_KWARGS, "phase": XXX, "metric_name": XXX)
Where "phase" (of Phase type) controls when to call torch.optim.lr_scheduler._LRScheduler.step().
For instance, in order to:
- Update LR on each batch: Use phase: Phase.TRAIN_BATCH_END
- Update LR after each epoch: Use phase: Phase.TRAIN_EPOCH_END
The "metric_name" refers to the metric to watch (See docs for "metric_to_watch" in train(...)
https://docs.deci.ai/super-gradients/docstring/training/sg_trainer.html) when using
ReduceLROnPlateau. In any other case this kwarg is ignored.
LR_SCHEDULER_KWARGS are simply passed to the torch scheduler's __init__.
For example:
lr_mode = {"StepLR": {"gamma": 0.1, "step_size": 1, "phase": Phase.TRAIN_EPOCH_END}}
is equivalent to following training code:
from torch.optim.lr_scheduler import StepLR
...
optimizer = ....
scheduler = StepLR(optimizer=optimizer, gamma=0.1, step_size=1)
for epoch in num_epochs:
train_epoch(...)
scheduler.step()
....
Examples
Using StepLRpythontrainer = Trainer("torch_Scheduler_example")The following code sections marked with '...' are placeholders
indicating additional necessary code that is not shown for simplicity.
train_dataloader = ...
valid_dataloader = ...
model = ...train_params = {
"max_epochs": 2,
"lr_mode": {"StepLR": {"gamma": 0.1, "step_size": 1, "phase": Phase.TRAIN_EPOCH_END}},
"lr_warmup_epochs": 0,
"initial_lr": 0.1,
"loss": torch.nn.CrossEntropyLoss(),
"optimizer": "SGD",
"criterion_params": {},
"optimizer_params": {"weight_decay": 1e-4, "momentum": 0.9},
"train_metrics_list": [Accuracy()],
"valid_metrics_list": [Accuracy()],
"metric_to_watch": "Accuracy",
"greater_metric_to_watch_is_better": True,
}
trainer.train(model=model, training_params=train_params, train_loader=dataloader, valid_loader=dataloader)
<details>
<summary>Equivalent in a <code>.yaml</code> configuration file:</summary>
yamltraining_hyperparams:
# Setting up LR Scheduler
lr_mode:
StepLR:
gamma: 0.1
step_size: 1
phase: TRAIN_EPOCH_END
# Setting up other parameters
max_epochs: 2
lr_warmup_epochs: 0
initial_lr: 0.1
loss: CrossEntropyLoss
optimizer: SGD
criterion_params: {}
optimizer_params:
weight_decay: 1e-4
momentum: 0.9
train_metrics_list:
- Accuracy
valid_metrics_list:
- Accuracy
metric_to_watch: Accuracy
greater_metric_to_watch_is_better: true...
</details>
Using
ReduceLROnPlateauIf you choose to use
ReduceLROnPlateau as the learning rate scheduler, you need to specify a metric_name.
This parameter follows the same guidelines as metric_to_watch. For an in-depth understanding of these metrics,
see the metrics guide.
pythontrainer = Trainer("torch_ROP_Scheduler_example")
train_dataloader = ...
valid_dataloader = ...
model = ...
train_params = {
"max_epochs": 2,
"lr_decay_factor": 0.1,
"lr_mode": {
"ReduceLROnPlateau": {"patience": 0, "phase": Phase.TRAIN_EPOCH_END, "metric_name": "DummyMetric"}},
"lr_warmup_epochs": 0,
"initial_lr": 0.1,
"loss": torch.nn.CrossEntropyLoss(),
"optimizer": "SGD",
"criterion_params": {},
"optimizer_params": {"weight_decay": 1e-4, "momentum": 0.9},
"train_metrics_list": [Accuracy()],
"valid_metrics_list": [Accuracy()],
"metric_to_watch": "DummyMetric",
"greater_metric_to_watch_is_better": True,
}
trainer.train(model=model, training_params=train_params, train_loader=dataloader, valid_loader=dataloader)
The scheduler's
state_dict is saved under torch_scheduler_state_dict entry inside the checkpoint during training,
allowing us to resume from the same state of the scheduling.<details>
<summary>Equivalent in a <code>.yaml</code> configuration file:</summary>
yamltraining_hyperparams:
# Setting up LR Scheduler
lr_mode:
ReduceLROnPlateau:
patience: 0
phase: TRAIN_EPOCH_END
metric_name: DummyMetric # Setting up other parameters
max_epochs: 2
lr_decay_factor: 0.1
lr_warmup_epochs: 0
initial_lr: 0.1
loss: CrossEntropyLoss
optimizer: SGD
criterion_params: {}
optimizer_params:
weight_decay: 1e-4
momentum: 0.9
train_metrics_list:
- Accuracy
valid_metrics_list:
- Accuracy
metric_to_watch: DummyMetric
greater_metric_to_watch_is_better: true
...
</details>---
Documentation/Source/Metrics (documentation/source/Metrics.md)
Metrics
The purpose of metrics is to allow you to monitor and quantify the training process. Therefore, metrics are an essential component in every deep learning training process.
For this purpose, we leverage the torchmetrics library.
From the
torchmetrics homepage: "TorchMetrics is a collection of 90+ PyTorch metrics implementations and an easy-to-use API to create custom metrics. It offers:
- A standardized interface to increase reproducibility
- Reduces Boilerplate
- Distributed-training compatible
- Rigorously tested
- Automatic accumulation over batches
- Automatic synchronization between multiple devices."
SG is compatible with any module metric implemented by torchmetrics (see complete list here).
Apart from the native
torchmetrics implementations, SG implements some metrics as torchmetrics.Metric objects as well: Accuracy
Top5
DetectionMetrics
IoU
PixelAccuracy
BinaryIOU
Dice
BinaryDice
DetectionMetrics_050
DetectionMetrics_075
DetectionMetrics_050_095
Basic Usage of Implemented Metrics
For coded training scripts (i.e., not using configuration files), the most basic usage is simply passing the metric objects through
train_metrics_list and valid_metrics_list:pythonfrom super_gradients import Trainer
...
from super_gradients.training.metrics import Accuracy, Top5trainer = Trainer("my_experiment")
train_dataloader = ...
valid_dataloader = ...
model = ...
train_params = {
...
"train_metrics_list": [Accuracy(), Top5()],
"valid_metrics_list": [Accuracy(), Top5()],
"metric_to_watch": "Accuracy",
"greater_metric_to_watch_is_better": True,
}
trainer.train(model=model, training_params=train_params, train_loader=train_dataloader, valid_loader=valid_dataloader)
Now, the metrics progress over the training epochs (and validation) will be displayed and logged in the Tensorboards, and any 3rd party SG Logger (see integration with Weights & Biases and Clearml in repo homepage).
Metric results will be lowercase, with the appropriate suffix:
train_accuracy, train_top5, valid_accuracy, valid_top5.
Also, notice the metric_to_watch set to Accuracy and greater_metric_to_watch_is_better=True, meaning that we will monitor the validation accuracy and save checkpoints according to it.
Open any of the tutorial notebooks to see the metrics monitoring in action.
For more info on checkpoints and logs, follow our SG checkpoints tutorial.Equivalently, for training with configuration files, your
my_training_hyperparams.yaml would contain:
yamldefaults:
- default_train_params
...
...
metric_to_watch: Accuracy
greater_metric_to_watch_is_better: Truetrain_metrics_list: # metrics for evaluation
- Accuracy
- Top5
valid_metrics_list: # metrics for evaluation
- Accuracy
- Top5
Using Custom Metrics
Suppose you implemented your own
MyAccuracy (more information on how to do so here), for coded training, you can pass an instance of it as done in the previous sub-section.
For training with configuration files, first decorate your metric class with SG's @register_metric decorator:
pythonfrom torchmetrics import Metric
import torch
from super_gradients.common.registry import register_metric@register_metric("my_accuracy")
class MyAccuracy(Metric):
def __init__(self):
super().__init__()
self.add_state("correct", default=torch.tensor(0), dist_reduce_fx="sum")
self.add_state("total", default=torch.tensor(0), dist_reduce_fx="sum")
def update(self, preds: torch.Tensor, target: torch.Tensor):
preds, target = self._input_format(preds, target)
assert preds.shape == target.shape
self.correct += torch.sum(preds == target)
self.total += target.numel()
def compute(self):
return self.correct.float() / self.total
Next, use the registered metric in your
my_training_hyperparams.yaml by plugging in the registered name, just as if it was any other metric:
yamldefaults:
- default_train_params
...
...
metric_to_watch: my_accuracy
greater_metric_to_watch_is_better: Truetrain_metrics_list: # metrics for evaluation
- my_accuracy
...
valid_metrics_list: # metrics for evaluation
- my_accuracy
...
Last, in your
my_train_from_recipe_script.py file, import the newly registered class (even though the class itself is unused, just to trigger the registry):
pythonfrom omegaconf import DictConfig
import hydra
import pkg_resources
from my_accuracy import MyAccuracy
from super_gradients import Trainer, init_trainer
@hydra.main(config_path=pkg_resources.resource_filename("super_gradients.recipes", ""), version_base="1.2")
def main(cfg: DictConfig) -> None:
Trainer.train_from_config(cfg)
def run():
init_trainer()
main()
if __name__ == "__main__":
run()
---
Documentation/Source/Model Zoo (documentation/source/model_zoo.md)
Model Zoo
Computer Vision Models - Pretrained Checkpoints
You can load any of our pretrained model in 2 lines of code:
pythonfrom super_gradients.training import models
from super_gradients.common.object_names import Modelsmodel = models.get(Models.YOLOX_S, pretrained_weights="coco")
All the available models are listed in the column
Model name.
Pretrained Classification PyTorch Checkpoints
| Model | Model name | Dataset | Resolution | Top-1 | Top-5 | Latency (HW)<sub>T4</sub> | Latency (Production)<sub>T4</sub> | Latency (HW)<sub>Jetson Xavier NX</sub> | Latency (Production)<sub>Jetson Xavier NX</sub> | Latency <sub>Cascade Lake</sub> | Torch Compile Support |
|-------------------------------|-----------------------|-------------|------------|--------|---------|----------------------------|-------------------------------------|------------------------------------------|---------------------------------------------------|:-------------------------------:|:---------------------:|
| ViT base | vit_base | ImageNet21K | 224x224 | 84.15 | - | 4.46ms | 4.60ms | - * | - | 57.22ms | Not Supported |
| ViT large | vit_large | ImageNet21K | 224x224 | 85.64 | - | 12.81ms | 13.19ms | - * | - | 187.22ms | Not Supported |
| BEiT | beit_base_patch16_224 | ImageNet21K | 224x224 | - | - | -ms | -ms | - * | - | -ms | Supported |
| EfficientNet B0 | efficientnet_b0 | ImageNet | 224x224 | 77.62 | 93.49 | 0.93ms | 1.38ms | - * | - | 3.44ms | Supported |
| RegNet Y200 | regnetY200 | ImageNet | 224x224 | 70.88 | 89.35 | 0.63ms | 1.08ms | 2.16ms | 2.47ms | 2.06ms | Supported |
| RegNet Y400 | regnetY400 | ImageNet | 224x224 | 74.74 | 91.46 | 0.80ms | 1.25ms | 2.62ms | 2.91ms | 2.87ms | Supported |
| RegNet Y600 | regnetY600 | ImageNet | 224x224 | 76.18 | 92.34 | 0.77ms | 1.22ms | 2.64ms | 2.93ms | 2.39ms | Supported |
| RegNet Y800 | regnetY800 | ImageNet | 224x224 | 77.07 | 93.26 | 0.74ms | 1.19ms | 2.77ms | 3.04ms | 2.81ms | Supported |
| ResNet 18 | resnet18 | ImageNet | 224x224 | 70.6 | 89.64 | 0.52ms | 0.95ms | 2.01ms | 2.30ms | 4.56ms | Supported |
| ResNet 34 | resnet34 | ImageNet | 224x224 | 74.13 | 91.7 | 0.92ms | 1.34ms | 3.57ms | 3.87ms | 7.64ms | Supported |
| ResNet 50 | resnet50 | ImageNet | 224x224 | 81.91 | 93.0 | 1.03ms | 1.44ms | 4.78ms | 5.10ms | 9.25ms | Supported |
| MobileNet V3_large-300 epochs | mobilenet_v3_large | ImageNet | 224x224 | 74.52 | 91.92 | 0.67ms | 1.11ms | 2.42ms | 2.71ms | 1.76ms | Supported |
| MobileNet V3_small | mobilenet_v3_small | ImageNet | 224x224 | 67.45 | 87.47 | 0.55ms | 0.96ms | 2.01ms * | 2.35ms | 1.06ms | Supported |
| MobileNet V2_w1 | mobilenet_v2 | ImageNet | 224x224 | 73.08 | 91.1 | 0.46 ms | 0.89ms | 1.65ms * | 1.90ms | 1.56ms | Supported |
NOTE: <br/>
- Latency (HW)* - Hardware performance (not including IO)<br/>
- Latency (Production) - Production Performance (including IO)
- Performance measured for T4 and Jetson Xavier NX with TensorRT, using FP16 precision and batch size 1
- Performance measured for Cascade Lake CPU with OpenVINO, using FP16 precision and batch size 1
Pretrained Object Detection PyTorch Checkpoints
| Model | Model Name | Dataset | Resolution | mAP<sup>val<br>0.5:0.95 | Latency (HW)<sub>T4</sub> | Latency (Production)<sub>T4</sub> | Latency (HW)<sub>Jetson Xavier NX</sub> | Latency (Production)<sub>Jetson Xavier NX</sub> | Latency <sub>Cascade Lake</sub> | Torch Compile Support |
|-----------------------|-----------------------|---------|------------|--------------------------|-------------------------------|-------------------------------------|------------------------------------------|---------------------------------------------------|:-------------------------------:|:---------------------:|
| YOLO-NAS S | yolo_nas_s | COCO | 640x640 | 47.5(FP16) 47.03(INT8) | 3.21(FP16) 2.36(INT8) | | | | | Supported |
| YOLO-NAS M | yolo_nas_m | COCO | 640x640 | 51.55(FP16) 51.0(INT8) | 5.85(FP16) 3.78(INT8) | | | | | Supported |
| YOLO-NAS L | yolo_nas_l | COCO | 640x640 | 52.22(FP16) 52.1(INT8) | 7.87(FP16) 4.78(INT8) | | | | | Supported |
| PP-YOLOE small | ppyoloe_s | COCO | 640x640 | 42.52 | 2.39ms | 4.3ms | 14.28ms | 14.99ms | - | Not Supported |
| PP-YOLOE medium | ppyoloe_m | COCO | 640x640 | 47.11 | 5.16ms | 7.05ms | 32.71ms | 33.46ms | - | Not Supported |
| PP-YOLOE large | ppyoloe_l | COCO | 640x640 | 49.48 | 7.65ms | 9.59ms | 51.13ms | 50.39ms | - | Not Supported |
| PP-YOLOE x-large | ppyoloe_x | COCO | 640x640 | 51.15 | 14.04ms | 15.96ms | 94.92ms | 94.22ms | - | Not Supported |
| YOLOX nano | yolox_n | COCO | 640x640 | 26.77 | 2.47ms | 4.09ms | 11.49ms | 12.97ms | - | Not Supported |
| YOLOX tiny | yolox_t | COCO | 640x640 | 37.18 | 3.16ms | 4.61ms | 15.23ms | 19.24ms | - | Not Supported |
| YOLOX small | yolox_s | COCO | 640x640 | 40.47 | 3.58ms | 4.94ms | 18.88ms | 22.48ms | - | Not Supported |
| YOLOX medium | yolox_m | COCO | 640x640 | 46.4 | 6.40ms | 7.65ms | 39.22ms | 44.5ms | - | Not Supported |
| YOLOX large | yolox_l | COCO | 640x640 | 49.25 | 10.07ms | 11.12ms | 68.73ms | 77.01ms | - | Not Supported |
| SSD lite MobileNet v2 | ssd_lite_mobilenet_v2 | COCO | 320x320 | 21.5 | 0.77ms | 1.40ms | 5.28ms | 6.44ms | 4.13ms | Not Supported |
| SSD lite MobileNet v1 | ssd_mobilenet_v1 | COCO | 320x320 | 24.3 | 1.55ms | 2.84ms | 8.07ms | 9.14ms | 22.76ms | Not Supported |
NOTE: <br/>
- Latency (HW)* - Hardware performance (not including IO)<br/>
- Latency (Production) - Production Performance (including IO)
- Latency performance measured for T4 and Jetson Xavier NX with TensorRT, using FP16 precision and batch size 1
- Latency performance measured for Cascade Lake CPU with OpenVINO, using FP16 precision and batch size 1
Pretrained Semantic Segmentation PyTorch Checkpoints
| Model | Model Name | Dataset | Resolution | mIoU | Latency b1<sub>T4</sub> | Latency b1<sub>T4</sub> including IO | Latency (Production)<sub>Jetson Xavier NX</sub> | Torch Compile Support |
|-----------------------|-------------------|------------|------------|-------|-------------------------|--------------------------------------|:-------------------------------------------------:|:---------------------:|
| PP-LiteSeg B50 | pp_lite_b_seg50 | Cityscapes | 512x1024 | 76.48 | 4.18ms | 31.22ms | 31.69ms | Supported |
| PP-LiteSeg B75 | pp_lite_b_seg75 | Cityscapes | 768x1536 | 78.52 | 6.84ms | 33.69ms | 49.89ms | Supported |
| PP-LiteSeg T50 | pp_lite_t_seg50 | Cityscapes | 512x1024 | 74.92 | 3.26ms | 30.33ms | 26.20ms | Supported |
| PP-LiteSeg T75 | pp_lite_t_seg75 | Cityscapes | 768x1536 | 77.56 | 5.20ms | 32.28ms | 38.03ms | Supported |
| DDRNet 23 slim | ddrnet_23_slim | Cityscapes | 1024x2048 | 79.41 | 5.74ms | 32.01ms | 45.18ms | Supported |
| DDRNet 23 | ddrnet_23 | Cityscapes | 1024x2048 | 81.48 | 12.74ms | 39.01ms | 106.26ms | Supported |
| DDRNet 39 | ddrnet_39 | Cityscapes | 1024x2048 | 81.32 | 23.57ms | 52.41ms | 145.79ms | Supported |
| STDC 1-Seg50 | stdc1_seg50 | Cityscapes | 512x1024 | 75.11 | 3.34ms | 30.12ms | 27.54ms | Supported |
| STDC 1-Seg75 | stdc1_seg75 | Cityscapes | 768x1536 | 77.8 | 5.53ms | 32.490ms | 43.88 | Supported |
| STDC 2-Seg50 | stdc2_seg50 | Cityscapes | 512x1024 | 76.44 | 4.12ms | 30.94ms | 32.03ms | Supported |
| STDC 2-Seg75 | stdc2_seg75 | Cityscapes | 768x1536 | 78.93 | 6.95ms | 33.89ms | 54.48ms | Supported |
| RegSeg (exp48) | regseg48 | Cityscapes | 1024x2048 | 78.15 | 12.03ms | 38.91ms | 78.20ms | Supported |
NOTE: <br/>
- Performance measured on T4 GPU with TensorRT, using FP16 precision and batch size 1 (latency), and not including IO
- For resolutions below 1024x2048 we first resize the input to the inference resolution and then resize the predictions to 1024x2048. The time of resizing is included in the measurements so that the practical input-size is 1024x2048.
- DDRNet23 and DDRNet23_Slim results were achieved with channel wise knowledge distillation training recipe.
Pretrained Pose Estimation PyTorch Checkpoints
| Model | Model Name | Dataset | Resolution | AP (No TTA / H-Flip TTA / H-Flip TTA+Rescoring) | Latency b1<sub>T4</sub> | Latency b1<sub>T4</sub> including IO | Latency (Production)<sub>Jetson Xavier NX</sub> |
|----------------|-----------------|-------------|------------|-------------------------------------------------|-------------------------|--------------------------------------|:-------------------------------------------------:|
| DEKR_W32_NO_DC | dekr_w32_no_dc | COCO2017 PE | 640x640 | 63.08 / 64.96 / 67.32 | 13.29 ms | 15.31 ms | 75.99 ms |
| YoloNAS POSE N | yolo_nas_pose_n | COCO2017 PE | 640x640 | 59.68 / N/A / N/A | N/A | 2.35 ms | 15.99 ms |
| YoloNAS POSE S | yolo_nas_pose_s | COCO2017 PE | 640x640 | 64.15 / N/A / N/A | N/A | 3.29 ms | 21.01 ms |
| YoloNAS POSE M | yolo_nas_pose_m | COCO2017 PE | 640x640 | 67.87 / N/A / N/A | N/A | 6.87 ms | 38.40 ms |
| YoloNAS POSE L | yolo_nas_pose_l | COCO2017 PE | 640x640 | 68.24 / N/A / N/A | N/A | 8.86 ms | 49.34 ms |
Implemented Model Architectures
Image Classification
- DensNet (Densely Connected Convolutional Networks) - Densely Connected Convolutional Networks https://arxiv.org/pdf/1608.06993.pdf
- DPN - Dual Path Networks https://arxiv.org/pdf/1707.01629
- EfficientNet - https://arxiv.org/abs/1905.11946
- GoogleNet - https://arxiv.org/pdf/1409.4842
- LeNet - https://yann.lecun.com/exdb/lenet/
- MobileNet - Efficient Convolutional Neural Networks for Mobile Vision Applications https://arxiv.org/pdf/1704.04861
- MobileNet v2 - https://arxiv.org/pdf/1801.04381
- MobileNet v3 - https://arxiv.org/pdf/1905.02244
- PNASNet - Progressive Neural Architecture Search Networks https://arxiv.org/pdf/1712.00559
- Pre-activation ResNet - https://arxiv.org/pdf/1603.05027
- RegNet - https://arxiv.org/pdf/2003.13678.pdf
- RepVGG - Making VGG-style ConvNets Great Again https://arxiv.org/pdf/2101.03697.pdf
- ResNet - Deep Residual Learning for Image Recognition https://arxiv.org/pdf/1512.03385
- ResNeXt - Aggregated Residual Transformations for Deep Neural Networks https://arxiv.org/pdf/1611.05431
- SENet - Squeeze-and-Excitation Networkshttps://arxiv.org/pdf/1709.01507
- ShuffleNet - https://arxiv.org/pdf/1707.01083
- ShuffleNet v2 - Efficient Convolutional Neural Network for Mobile
Deviceshttps://arxiv.org/pdf/1807.11164
- VGG - Very Deep Convolutional Networks for Large-scale Image Recognition https://arxiv.org/pdf/1409.1556
Object Detection
- CSP DarkNet
- DarkNet-53
- SSD (Single Shot Detector) - https://arxiv.org/pdf/1512.02325
- YOLOX - https://arxiv.org/abs/2107.08430
- PP-YoloE - https://arxiv.org/abs/2203.16250Semantic Segmentation
- PP-LiteSeg - https://arxiv.org/pdf/2204.02681v1.pdf
- DDRNet (Deep Dual-resolution Networks) - https://arxiv.org/pdf/2101.06085.pdf
- LadderNet - Multi-path networks based on U-Net for medical image segmentation https://arxiv.org/pdf/1810.07810
- RegSeg - Rethink Dilated Convolution for Real-time Semantic Segmentation https://arxiv.org/pdf/2111.09957
- ShelfNet - https://arxiv.org/pdf/1811.11254
- STDC - Rethinking BiSeNet For Real-time Semantic Segmentation https://arxiv.org/pdf/2104.13188
Pose Estimation
- HRNet DEKR - Bottom-Up Human Pose Estimation Via Disentangled Keypoint Regression https://arxiv.org/pdf/2104.02300.pdf
- YoloNAS Pose
---
Documentation/Source/ModelPredictions (documentation/source/ModelPredictions.md)
Using Pretrained Models for Predictions
In this tutorial, we will demonstrate how to use the
model.predict() method for object detection tasks.The model used in this tutorial is YOLO-NAS, pre-trained on the COCO dataset, which contains 80 object categories.
Warning: If you trained your model on a dataset that does not inherit from any of the SuperGradients dataset, you will need to follow some additional steps before running the model. You can find these steps in the following tutorial.
Note that the
model.predict() method is currently only available for detection tasks.Supported Media Formats
A
mode.predict() method is built to handle multiple data formats and types.
Here is the full list of what predict() method can handle:| Argument Semantics | Argument Type | Supported layout | Example | Notes |
|------------------------------------|--------------------|-----------------------------------|------------------------------------------------------------------------------------------------|--------------------------------------------------------------------------------------------------|
| Path to local image |
str | - | predict("path/to/image.jpg") | All common image extensions are supported. |
| Path to images directory | str | - | predict("path/to/images/directory") | |
| Path to local video | str | - | predict("path/to/video.mp4") | All common video extensions are supported. |
| URL to remote image | str | - | predict("https://example.com/image.jpg") | |
| 3-dimensional Numpy image | np.ndarray | [H, W, C] | predict(np.zeros((480, 640, 3), dtype=np.uint8)) | Channels last, RGB channel order for 3-channel images |
| 4-dimensional Numpy image | np.ndarray | [N, H, W, C] or [N, C, H, W] | predict(np.zeros((480, 640, 3), dtype=np.uint8)) | Tensor layout (NHWC or NCHW) is inferred w.r.t to number of input channels of underlying model |
| List of 3-dimensional numpy arrays | List[np.ndarray] | [H1, W1, C], [H2, W2, C], ... | predict([np.zeros((480, 640, 3), dtype=np.uint8), np.zeros((384, 512, 3), dtype=np.uint8) ]) | Images may vary in size, but should have same number of channels |
| 3-dimensional Torch Tensor | torch.Tensor | [H, W, C] or [C, H, W] | predict(torch.zeros((480, 640, 3), dtype=torch.uint8)) | Tensor layout (HWC or CHW) is inferred w.r.t to number of input channels of underlying model |
| 4-dimensional Torch Tensor | torch.Tensor | [N, H, W, C] or [N, C, H, W] | predict(torch.zeros((4, 480, 640, 3), dtype=torch.uint8)) | Tensor layout (NHWC or NCHW) is inferred w.r.t to number of input channels of underlying model |Important note - When using batched input (4-dimensional
np.ndarray or torch.Tensor) formats, normalization and size preprocessing will be applied to these inputs.
This means that the input tensors should not be normalized beforehand.
Here is the example of incorrect code of using model.predict():pythonIncorrect code example. Do not use it.
from super_gradients.training import dataloaders
from super_gradients.common.object_names import Models
from super_gradients.training import modelsval_loader = dataloaders.get("coco2017_val_yolo_nas")
model = models.get(Models.YOLO_NAS_L, pretrained_weights="coco")
for (inputs, *_) in val_loader: # Error here: inputs as already normalized by dataset class
model.predict(inputs).show() # This will not work as expected
Since
model.predict() encapsulates normalization and size preprocessing, it is not designed to handle pre-normalized images as input.
Please keep this in mind when using model.predict() with batched inputs.
Detect Objects in Multiple Images
#### Load the Model and Prepare the Images
First, let's load the pre-trained
Yolo-NAS model using the models.get() function and define a list of image paths or URLs that we want to process:pythonfrom super_gradients.common.object_names import Models
from super_gradients.training import modelsmodel = models.get(Models.YOLO_NAS_L, pretrained_weights="coco")
Detect Objects in the Images
The model.predict() method returns an ImagesDetectionPrediction object, which contains the detection results for each image.pythonIMAGES = [
"path/to/local/image1.jpg",
"path/to/local/image2.jpg",
"https://example.com/image3.jpg",
]images_predictions = model.predict(IMAGES)
You can use the default IoU and Confidence threshold or override them like this:pythonimages_predictions = model.predict(IMAGES, iou=0.5, conf=0.7)
- iou: IoU threshold for the non-maximum suppression (NMS) algorithm. If None, the default value associated with the model used.
- conf: Confidence threshold. Predictions below this threshold are discarded. If None, the default value associated with the model used.Display the Detected Objects
To display the detected objects and their bounding boxes on the images, call images_predictions.show().pythonimages_predictions.show()
You can customize the following optional parameters:
pythonimages_predictions.show(box_thickness=2, show_confidence=True)
- box_thickness: Thickness of bounding boxes.
- show_confidence: Whether to show confidence scores on the image.
- color_mapping: List of tuples representing the colors for each class.
- class_names: List of class names to display. Only classes that the model was trained on are supported. By default, show all these classes.Save the Images with Detected Objects
To save the images with detected objects as separate files, call the images_predictions.save() method and specify the output folder.
pythonimages_predictions.save(output_folder="output_folder/")
You can also customize the same parameters as in the
images_predictions.show() method:
pythonimages_predictions.save(output_folder="output_folder/", box_thickness=2, show_confidence=True)
Access Detection Results
To access the detection results for each image, you can iterate over the images_predictions object. For each detected object, you can retrieve various attributes such as the label ID, label name, confidence score, and bounding box coordinates. These attributes can be used for further processing or analysis.pythonfor image_prediction in images_predictions:
class_names = image_prediction.class_names
labels = image_prediction.prediction.labels
confidence = image_prediction.prediction.confidence
bboxes = image_prediction.prediction.bboxes_xyxy for i, (label, conf, bbox) in enumerate(zip(labels, confidence, bboxes)):
print("prediction: ", i)
print("label_id: ", label)
print("label_name: ", class_names[int(label)])
print("confidence: ", conf)
print("bbox: ", bbox)
print("--" * 10)
# You can use the detection results for various tasks, such as:
# - Filtering objects based on confidence scores or labels
# - Analyzing object distributions within the images
# - Calculating object dimensions or areas
# - Implementing custom visualization techniques
# - ...
You can use these detection results to implement any feature not implemented by SuperGradients to fit your specific needs.
You can also directly access a specific image prediction by referencing its index.
images_predictions[1] will give you the prediction of the second image.
Detect Objects in Animated GIFs and Videos
The processing for both gif and videos is similar, as they are treated as videos internally. You can use the same model.predict() method as before, but pass the path to a GIF or video file instead. The results can be saved as either a .gif or .mp4.To mitigate Out-of-Memory (OOM) errors, the
model.predict() method for video returns a generator object. This allows the video frames to be processed sequentially, minimizing memory usage. It's important to be aware that model inference in this mode will be slower since batching is not supported.Consequently, you need to invoke
model.predict() before each show() and save() call.Load an Animated GIF or Video File
Let's load an animated GIF or a video file and pass it to the model.predict() method:pythonMEDIA_PATH = "path/to/animated_gif_or_video.gif_or_mp4"
media_predictions = model.predict(MEDIA_PATH)
Display the Detected Objects
To display the detected objects and their bounding boxes in the animated GIF or video, call media_predictions.show():
pythonmedia_predictions.show()
Save the Results with Detected Objects
To save the results with detected objects as a separate file, call the media_predictions.save() method, and simply specify the desired output extension in the output name: .gif or .mp4Save as a
.gif
pythonmedia_predictions.save("output_video.gif") # Save as .gif
Save as a
.mp4
pythonmedia_predictions.save("output_video.mp4") # Save as .mp4
Frames Per Second (FPS)
The number of Frames Per Second (FPS) at which the model processes the gif/video can be seen directly next to the loading bar when running model.predict('my_video.mp4').In the following example, the FPS is 39.49it/s (i.e. fps)
Processing Video: 100%|βββββββββββββββββββββββ| 306/306 [00:07<00:00, 39.49it/s]Note that the video/gif will be saved with original FPS (i.e.
media_predictions.fps).Access Frame-by-Frame Detection Results for GIFs and Videos
Iterating over the media_predictions object allows you to access the detection results for each frame. This provides an opportunity to perform frame-specific operations, like applying custom filters or visualizations.pythonfor frame_index, frame_prediction in enumerate(media_predictions):
labels = frame_prediction.prediction.labels
confidence = frame_prediction.prediction.confidence
bboxes = frame_prediction.prediction.bboxes_xyxy # You can do any frame-specific operations
# ...
# Example: Save individual frames with detected objects
frame_name = f"output/frame_{frame_index}.jpg"
frame_prediction.save(frame_name) # save frame as an image
Detect Objects Using a Webcam
Call the model.predict_webcam() method to start detecting objects using your webcam:pythonmodel.predict_webcam()
The detected objects and their bounding boxes will be displayed on the webcam feed in real-time. Press 'q' to quit the webcam feed.
Note that
model.predict_webcam() and model.predict() share the same parameters.Frames Per Second (FPS)
In the case of a Webcam, contrary to when processing a video by batch, the number of Frames Per Seconds (FPS) directly affects the display FPS since we show each frame right after it is processed.You can find this information directly written in a corner of the video.
Using GPU for Object Detection
If your system has a GPU available, you can use it for faster object detection by moving the model to the GPU:pythonmodel = model.to("cuda" if torch.cuda.is_available() else "cpu")
model.predict(...)
This allows the model to run on the GPU, significantly speeding up the object detection process. Note that using a GPU requires having the necessary drivers and compatible hardware installed.
Skipping Image Resizing
Skipping image resizing in object detection can have a significant impact on the results. Typically, models are trained on images of a certain size, with (640, 640) being a common dimension.By default, the
model.predict(...) method resizes input images to the training size. However, there's an option to bypass this resizing step, which offers several benefits:- Speed Improvement for Smaller Images: If your original image is smaller than the typical training size, avoiding resizing can speed up the prediction process.
- Enhanced Detection of Small Objects in High-Resolution Images: For high-resolution images containing numerous small objects, processing the images in their original size can improve the model's ability to recall these objects. This comes at the expense of speed but can be beneficial for detailed analysis.
To apply this approach, simply use the
skip_image_resizing parameter in the model.predict(...) method as shown below:pythonpredictions = model.predict(image, skip_image_resizing=True)
#### Example
The following images illustrate the difference in detection results with and without resizing.
#### Original Image
This is the raw image before any processing.
#### Image Processed with Standard Resizing (640x640)
This image shows the detection results after resizing the image to the model's trained size of 640x640.
#### Image Processed in Original Size
Here, the image is processed in its original size, demonstrating how the model performs without resizing. Notice the differences in object detection and details compared to the resized version.
---
Documentation/Source/Models (documentation/source/models.md)
Models
SuperGradients provides an extensive collection of state-of-the-art (SOTA) models in its model zoo.
These models are implemented as
torch.nn.Module and can be used, customized, and trained like any other torch module.The 3 main use cases of the Model Zoo are to
- Train a model from scratch
- Fine-tune a pre-trained model
- Use a model (pre-trained or not) as the backbone of a larger architecture.
Instantiating a model
To instantiate a model, specify the model name and the number of classes desired.
pythonfrom super_gradients.training import modelsInstantiate resnet18 with head supporting 100 classes
default_resnet18 = models.get(model_name="resnet18", num_classes=100)
All model names are available in the model zoo,but can also be dynamically accessed through
super_gradients.common.object_names for autocompletion
pythonfrom super_gradients.training import models
from super_gradients.common import object_namesinstantiate default pretrained resnet18
default_resnet18 = models.get(model_name=object_names.Models.RESNET18, num_classes=100)
Instantiating a pretrained model
When loading a pre-trained model, SuperGradients also provides a pre-trained head by default.
The head's dimension is determined by the number of classes in the dataset used for training.If you're using a different dataset, you'll need to change the number of classes in the head.
This keeps all the pre-trained weights of the model intact, except for the head which will be new and untrained.
The model will not be able to predict accurately until fine-tuned.
With pretrained head
pythonfrom super_gradients.training import modelsWill reproduce the model zoo metrics on imagenet
model = models.get(model_name="resnet18", pretrained_weights="imagenet")
You can find the datasets used for pretraining our models in the model zoo, and specify it in the pretrained_weights.With new head
pythonfrom super_gradients.training import modelsCan be trained on a dataset of 94 classes
model = models.get(model_name="resnet18", num_classes=94, pretrained_weights="imagenet")
Loading a Backbone
In deep learning, a backbone is a pre-trained neural network that serves as a starting point to build a larger architecture.
It is typically a feature extractor trained on a large dataset and meant to capture important features of the data. When loading a model as a backbone in SuperGradients, you will get the model without the global pooling stage and the classifier head.
pythonfrom super_gradients.training import modelsinstantiate pretrained resnet18, without classifier head. Output will be from the last stage before global pooling
backbone_resnet18 = models.get(model_name="resnet18", arch_params={"backbone_mode": True}, pretrained_weights="imagenet")
This backbone model can later be used as part of another model
pythonimport torchclass CustomModel(torch.nn.Module):
def __init__(self, backbone):
super().__init__()
self._backbone = backbone
self._head = ...
def forward(self, x):
out = self._backbone(x)
out = self._head(out)
return out
model = CustomModel(backbone=backbone_resnet18)
Playing with the model architecture parameters
All of SuperGradients model architectures can be parametrized using
arch_params.
You can find the documentation about parameters of every architecture, and their default values, in the recipes.
In this example, we override the default params of efficientnet_b0
pythonfrom super_gradients.training import modelsarch_params = {
"drop_connect_rate": 0.3,
"image_size": 500,
}
yolox_custom = models.get(model_name="efficientnet_b0", arch_params=arch_params, num_classes=15)
---
Documentation/Source/Models Export (documentation/source/models_export.md)
This tutorial shows how to export SG models to ONNX format for deployment to ONNX-compatible runtimes and accelerators.
From this tutorial you will learn:
* How to export Object Detection model to ONNX and it with ONNXRuntime / TensorRT
* How to enable FP16 / INT8 quantization and export a model with calibration
* How to customize NMS parameters and number of detections per image
* How to choose whether to use TensorRT or ONNXRuntime as a backend
New Export API
A new export API is introduced in SG 3.2.0. It is aimed to simplify the export process and allow end-to-end export of SG models to ONNX format with a single line of code.
Currently supported models
- YoloNAS
- PPYoloE
Supported features
- Exporting a model to OnnxRuntime and TensorRT
- Exporting a model with preprocessing (e.g. normalizing/standardizing image according to normalization parameters during training)
- Exporting a model with postprocessing (e.g. predictions decoding and NMS) - you obtain the ready-to-consume bounding box outputs
- FP16 / INT8 quantization support with calibration
- Pre- and post-processing steps can be customized by the user if needed
- Customising input image shape and batch size
- Customising NMS parameters and number of detections per image
- Customising output format (flat or batched)
python!pip install -qq super_gradients==3.4.0
Minimalistic export example
Let start with the most simple example of exporting a model to ONNX format.
We will use YoloNAS-S model in this example. All models that suports new export API now expose a
export() method that can be used to export a model. There is one mandatory argument that should be passed to the export() method - the path to the output file. Currently, only .onnx format is supported, but we may add support for CoreML and other formats in the future.
pythonfrom super_gradients.common.object_names import Models
from super_gradients.training import modelsmodel = models.get(Models.YOLO_NAS_S, pretrained_weights="coco")
export_result = model.export("yolo_nas_s.onnx")
A lot of work just happened under the hood:
* A model was exported to ONNX format using default batch size of 1 and input image shape that was used during training
* A preprocessing and postprocessing steps were attached to ONNX graph
* For pre-processing step, the normalization parameters were extracted from the model itself (to be consistent with the image normalization and channel order used during training)
* For post-processing step, the NMS parameters were also extracted from the model and NMS module was attached to the graph
* ONNX graph was checked and simplified to improve compatibility with ONNX runtimes.
A returned value of
export() method is an instance of ModelExportResult class.
First of all it serves the purpose of storing all the information about the exported model in a single place.
It also provides a convenient way to get an example of running the model and getting the output:
pythonexport_result
Model exported successfully to yolo_nas_s.onnx
Model expects input image of shape [1, 3, 640, 640]
Input image dtype is torch.uint8
Exported model already contains preprocessing (normalization) step, so you don't need to do it manually.
Preprocessing steps to be applied to input image are:
Sequential(
(0): CastTensorTo(dtype=torch.float32)
(1): ApplyMeanStd(mean=[0.], scale=[255.])
)
Exported model contains postprocessing (NMS) step with the following parameters:
num_pre_nms_predictions=1000
max_predictions_per_image=1000
nms_threshold=0.7
confidence_threshold=0.25
output_predictions_format=batch
Exported model is in ONNX format and can be used with ONNXRuntime
To run inference with ONNXRuntime, please use the following code snippet:
import onnxruntime
import numpy as np
session = onnxruntime.InferenceSession("yolo_nas_s.onnx", providers=["CUDAExecutionProvider", "CPUExecutionProvider"])
inputs = [o.name for o in session.get_inputs()]
outputs = [o.name for o in session.get_outputs()]
example_input_image = np.zeros((1, 3, 640, 640)).astype(np.uint8)
predictions = session.run(outputs, {inputs[0]: example_input_image})
Exported model has predictions in batch format:
num_detections, pred_boxes, pred_scores, pred_classes = predictions
for image_index in range(num_detections.shape[0]):
for i in range(num_detections[image_index,0]):
class_id = pred_classes[image_index, i]
confidence = pred_scores[image_index, i]
x_min, y_min, x_max, y_max = pred_boxes[image_index, i]
print(f"Detected object with class_id={class_id}, confidence={confidence}, x_min={x_min}, y_min={y_min}, x_max={x_max}, y_max={y_max}")
That's it. You can now use the exported model with any ONNX-compatible runtime or accelerator.
pythonimport cv2
import numpy as np
from super_gradients.training.utils.media.image import load_image
import onnxruntimeimage = load_image("https://deci-pretrained-models.s3.amazonaws.com/sample_images/beatles-abbeyroad.jpg")
image = cv2.resize(image, (export_result.input_image_shape[1], export_result.input_image_shape[0]))
image_bchw = np.transpose(np.expand_dims(image, 0), (0, 3, 1, 2))
session = onnxruntime.InferenceSession(export_result.output, providers=["CUDAExecutionProvider", "CPUExecutionProvider"])
inputs = [o.name for o in session.get_inputs()]
outputs = [o.name for o in session.get_outputs()]
result = session.run(outputs, {inputs[0]: image_bchw})
result[0].shape, result[1].shape, result[2].shape, result[3].shape
((1, 1), (1, 1000, 4), (1, 1000), (1, 1000))
In the next section we unpack the result of prediction and show how to use it.
Output format for detection models
If
preprocessing=True (default value) then all models will be exported with NMS. If preprocessing=False models will be exported without NMS and raw model outputs will be returned. In this case, you will need to apply NMS yourself. This is useful if you want to use a custom NMS implementation that is not ONNX-compatible. In most cases you will want to use default preprocessing=True. It is also possible to pass a custom nn.Module as a postprocessing argument to the export() method. This module will be attached to the exported ONNX graph instead of the default NMS module. We encourage users to read the documentation of the export() method to learn more about the advanced options.When exporting an object detection model with postprocessing enabled, the prediction format can be one of two:
* A "flat" format -
DetectionOutputFormatMode.FLAT_FORMAT
* A "batched" format - DetectionOutputFormatMode.BATCH_FORMATYou can select the desired output format by setting
export(..., output_predictions_format=DetectionOutputFormatMode.BATCH_FORMAT).Flat format
A detection results returned as a single tensor of shape
[N, 7], where N is the number of detected objects in the entire batch. Each row in the tensor represents a single detection result and has the following format:[batch_index, x1, y1, x2, y2, class score, class index]When exporting a model with batch size of 1 (default mode) you can ignore the first column as all boxes will belong to the single sample. In case you export model with batch size > 1 you have to iterate over this array like so:
pythonfor sample_index in export_result.batch_size:
detections_for_sample_i = flat_predictions[flat_predictions[:, 0] == sample_index]
for (x1, y1, x2, y2, class_score, class_index) in detections_for_sample_i:
class_index = int(class_index) # convert from float to int
# do something with the detection predictions
Batch format
A second supported format is so-called "batch". It matches with output format of TensorRT's NMS implementation. The return value in this case is tuple of 4 tensors:
*
num_predictions - [B, 1] - A number of predictions per sample
* pred_boxes - [B, N, 4] - A coordinates of the predicted boxes in X1, Y1, X2, Y2 format
* pred_scores - [B, N] - A scores of the predicted boxes
* pred_classes - [B, N] - A class indices of the predicted boxesHere
B corresponds to batch size and N is the maximum number of detected objects per image.
In order to get the actual number of detections per image you need to iterate over num_predictions tensor and get the first element of each row.Now when you're familiar with the output formats, let's see how to use them.
To start, it's useful to take a look at the values of the predictions with a naked eye:
pythonnum_predictions, pred_boxes, pred_scores, pred_classes = result
num_predictions
array([[25]], dtype=int64)
pythonnp.set_printoptions(threshold=50, edgeitems=3)
pred_boxes, pred_boxes.shape
(array([[[439.55383, 253.22733, 577.5956 , 548.11975],
[ 35.71795, 249.40926, 176.62216, 544.69794],
[182.39618, 249.49301, 301.44122, 529.3324 ],
...,
[ -1. , -1. , -1. , -1. ],
[ -1. , -1. , -1. , -1. ],
[ -1. , -1. , -1. , -1. ]]], dtype=float32),
(1, 1000, 4))
pythonnp.set_printoptions(threshold=50, edgeitems=5)
pred_scores, pred_scores.shape
(array([[ 0.9694027, 0.9693378, 0.9665707, 0.9619047, 0.7538769, ...,
-1. , -1. , -1. , -1. , -1. ]],
dtype=float32),
(1, 1000))
pythonnp.set_printoptions(threshold=50, edgeitems=10)
pred_classes, pred_classes.shape
(array([[ 0, 0, 0, 0, 0, 0, 0, 0, 2, 2, ..., -1, -1, -1, -1, -1,
-1, -1, -1, -1, -1]], dtype=int64),
(1, 1000))
Visualizing predictions
For sake of this tutorial we will use a simple visualization function that is tailored for batch_size=1 only.
You can use it as a starting point for your own visualization code.
pythonfrom super_gradients.training.datasets.datasets_conf import COCO_DETECTION_CLASSES_LIST
from super_gradients.training.utils.detection_utils import DetectionVisualization
import matplotlib.pyplot as plt
def show_predictions_from_batch_format(image, predictions):
num_predictions, pred_boxes, pred_scores, pred_classes = predictions
assert num_predictions.shape[0] == 1, "Only batch size of 1 is supported by this function"
num_predictions = int(num_predictions.item())
pred_boxes = pred_boxes[0, :num_predictions]
pred_scores = pred_scores[0, :num_predictions]
pred_classes = pred_classes[0, :num_predictions]
image = image.copy()
class_names = COCO_DETECTION_CLASSES_LIST
color_mapping = DetectionVisualization._generate_color_mapping(len(class_names))
for (x1, y1, x2, y2, class_score, class_index) in zip(pred_boxes[:, 0], pred_boxes[:, 1], pred_boxes[:, 2], pred_boxes[:, 3], pred_scores, pred_classes):
image = DetectionVisualization.draw_box_title(
image_np=image,
x1=int(x1),
y1=int(y1),
x2=int(x2),
y2=int(y2),
class_id=class_index,
class_names=class_names,
color_mapping=color_mapping,
box_thickness=2,
pred_conf=class_score,
)
plt.figure(figsize=(8, 8))
plt.imshow(image)
plt.tight_layout()
plt.show()
pythonshow_predictions_from_batch_format(image, result)
Changing the output format
You can explicitly specify output format of the predictions by setting the
output_predictions_format argument of export() method. Let's see how it works:pythonfrom super_gradients.conversion import DetectionOutputFormatModeexport_result = model.export("yolo_nas_s.onnx", output_predictions_format=DetectionOutputFormatMode.FLAT_FORMAT)
export_result
Model exported successfully to yolo_nas_s.onnx
Model expects input image of shape [1, 3, 640, 640]
Input image dtype is torch.uint8
Exported model already contains preprocessing (normalization) step, so you don't need to do it manually.
Preprocessing steps to be applied to input image are:
Sequential(
(0): CastTensorTo(dtype=torch.float32)
(1): ApplyMeanStd(mean=[0.], scale=[255.])
)
Exported model contains postprocessing (NMS) step with the following parameters:
num_pre_nms_predictions=1000
max_predictions_per_image=1000
nms_threshold=0.7
confidence_threshold=0.25
output_predictions_format=flat
Exported model is in ONNX format and can be used with ONNXRuntime
To run inference with ONNXRuntime, please use the following code snippet:
import onnxruntime
import numpy as np
session = onnxruntime.InferenceSession("yolo_nas_s.onnx", providers=["CUDAExecutionProvider", "CPUExecutionProvider"])
inputs = [o.name for o in session.get_inputs()]
outputs = [o.name for o in session.get_outputs()]
example_input_image = np.zeros((1, 3, 640, 640)).astype(np.uint8)
predictions = session.run(outputs, {inputs[0]: example_input_image})
Exported model has predictions in flat format:
# flat_predictions is a 2D array of [N,7] shape
# Each row represents (image_index, x_min, y_min, x_max, y_max, confidence, class_id)
# Please note all values are floats, so you have to convert them to integers if needed
[flat_predictions] = predictions
for (_, x_min, y_min, x_max, y_max, confidence, class_id) in flat_predictions[0]:
class_id = int(class_id)
print(f"Detected object with class_id={class_id}, confidence={confidence}, x_min={x_min}, y_min={y_min}, x_max={x_max}, y_max={y_max}")
Now we exported a model that produces predictions in
flat format. Let's run the model like before and see the result:
pythonsession = onnxruntime.InferenceSession(export_result.output, providers=["CUDAExecutionProvider", "CPUExecutionProvider"])
inputs = [o.name for o in session.get_inputs()]
outputs = [o.name for o in session.get_outputs()]
result = session.run(outputs, {inputs[0]: image_bchw})
result[0].shape
(25, 7)
pythondef show_predictions_from_flat_format(image, predictions):
[flat_predictions] = predictions image = image.copy()
class_names = COCO_DETECTION_CLASSES_LIST
color_mapping = DetectionVisualization._generate_color_mapping(len(class_names))
for (sample_index, x1, y1, x2, y2, class_score, class_index) in flat_predictions[flat_predictions[:, 0] == 0]:
class_index = int(class_index)
image = DetectionVisualization.draw_box_title(
image_np=image,
x1=int(x1),
y1=int(y1),
x2=int(x2),
y2=int(y2),
class_id=class_index,
class_names=class_names,
color_mapping=color_mapping,
box_thickness=2,
pred_conf=class_score,
)
plt.figure(figsize=(8, 8))
plt.imshow(image)
plt.tight_layout()
plt.show()
pythonshow_predictions_from_flat_format(image, result)
Changing postprocessing settings
You can control a number of parameters in the NMS settings as well as maximum number of detections per image before and after NMS step:
* IOU threshold for NMS -
nms_iou_threshold
* Score threshold for NMS - nms_score_threshold
* Maximum number of detections per image before NMS - max_detections_before_nms
* Maximum number of detections per image after NMS - max_detections_after_nmsFor sake of demonstration, let's export a model that would produce at most one detection per image with confidence threshold above 0.8 and NMS IOU threshold of 0.5. Let's use at most 100 predictions per image before NMS step:
pythonexport_result = model.export(
"yolo_nas_s_top_1.onnx",
confidence_threshold = 0.8,
nms_threshold = 0.5,
num_pre_nms_predictions = 100,
max_predictions_per_image = 1,
output_predictions_format = DetectionOutputFormatMode.FLAT_FORMAT
)session = onnxruntime.InferenceSession(export_result.output, providers=["CUDAExecutionProvider", "CPUExecutionProvider"])
inputs = [o.name for o in session.get_inputs()]
outputs = [o.name for o in session.get_outputs()]
result = session.run(outputs, {inputs[0]: image_bchw})
show_predictions_from_flat_format(image, result)
Export of quantized model
You can export a model with quantization to FP16 or INT8. To do so, you need to specify the
quantization_mode argument of export() method.Important notes:
* Quantization to FP16 requires CUDA / MPS device available and would not work on CPU-only machines.
Let's see how it works:
pythonfrom super_gradients.conversion.conversion_enums import ExportQuantizationModeexport_result = model.export(
"yolo_nas_s_int8.onnx",
output_predictions_format=DetectionOutputFormatMode.FLAT_FORMAT,
quantization_mode=ExportQuantizationMode.INT8 # or ExportQuantizationMode.FP16
)
session = onnxruntime.InferenceSession(export_result.output, providers=["CUDAExecutionProvider", "CPUExecutionProvider"])
inputs = [o.name for o in session.get_inputs()]
outputs = [o.name for o in session.get_outputs()]
result = session.run(outputs, {inputs[0]: image_bchw})
show_predictions_from_flat_format(image, result)
Advanced INT-8 quantization options
When quantizing a model using
quantization_mode==ExportQuantizationMode.INT8 you can pass a DataLoader to export() function to collect correct statistics of activations to prodice a more accurate quantized model.
We expect the DataLoader to return either a tuple of tensors or a single tensor. In case a tuple of tensors is returned by data-loader the first element will be used as input image.
You can use existing data-loaders from SG here as is.Important notes
* A
calibration_loader should use same image normalization parameters that were used during training.In the example below we use a dummy data-loader for sake of showing how to use this feature. You should use your own data-loader here.
pythonimport torch
from torch.utils.data import DataLoader
from super_gradients.conversion import ExportQuantizationModeTHIS IS ONLY AN EXAMPLE. YOU SHOULD USE YOUR OWN DATA-LOADER HERE
dummy_calibration_dataset = [torch.randn((3, 640, 640), dtype=torch.float32) for _ in range(32)]
dummy_calibration_loader = DataLoader(dummy_calibration_dataset, batch_size=8, num_workers=0)
THIS IS ONLY AN EXAMPLE. YOU SHOULD USE YOUR OWN DATA-LOADER HERE
export_result = model.export(
"yolo_nas_s_int8_with_calibration.onnx",
output_predictions_format=DetectionOutputFormatMode.FLAT_FORMAT,
quantization_mode=ExportQuantizationMode.INT8,
calibration_loader=dummy_calibration_loader
)
session = onnxruntime.InferenceSession(export_result.output, providers=["CUDAExecutionProvider", "CPUExecutionProvider"])
inputs = [o.name for o in session.get_inputs()]
outputs = [o.name for o in session.get_outputs()]
result = session.run(outputs, {inputs[0]: image_bchw})
show_predictions_from_flat_format(image, result)
25%|βββββββββββββββββββββββββββββββββ | 4/16 [00:11<00:34, 2.91s/it]
Limitations
* Dynamic batch size / input image shape is not supported yet. You can only export a model with a fixed batch size and input image shape.
* TensorRT of version 8.4.1 or higher is required.
* Quantization to FP16 requires CUDA / MPS device available.
Supported backends
Currently, we support two backends for exporting models:
* ONNX Runtime
* TensorRT
The only difference between these two backends is what NMS implementation will be used.
ONNX Runtime uses NMS implementation from ONNX opset, while TensorRT uses its own NMS implementation which is expected to be faster.
A disadvantage of TensorRT backend is that you cannot run model exported for TensorRT backend by ONNX Runtime.
You can, however, run models exported for ONNX Runtime backend inside TensorRT.
Therefore, ONNX Runtime backend is recommended for most use-cases and is used by default.
You can specify the desired execution backend by setting the
execution_backend argument of export() method:pythonfrom super_gradients.conversion import ExportTargetBackendmodel.export(..., engine=ExportTargetBackend.ONNXRUNTIME)
pythonfrom super_gradients.conversion import ExportTargetBackendmodel.export(..., engine=ExportTargetBackend.TENSORRT)
Legacy low-level export API
The .export() API is a new high-level API that is recommended for most use-cases.
However old low-level API is still available for advanced users:
* https://docs.deci.ai/super-gradients/docstring/training/models.html#training.models.conversion.convert_to_onnx
* https://docs.deci.ai/super-gradients/docstring/training/models.html#training.models.conversion.convert_to_coreml
---
Documentation/Source/Models Export Pose (documentation/source/models_export_pose.md)
Pose Estimation Models Export
This tutorial shows how to export YoloNAS-Pose model to ONNX format for deployment to ONNX-compatible runtimes and accelerators.
From this tutorial you will learn:
* How to export YoloNAS-Pose model to ONNX and run it with ONNXRuntime / TensorRT
* How to enable FP16 / INT8 quantization and export a model with calibration
* How to customize NMS parameters and number of detections per image
* How to choose whether to use TensorRT or ONNXRuntime as a backend
Supported pose estimation models
- YoloNAS-Pose N,S,M,L
Supported features
- Exporting a model to OnnxRuntime and TensorRT
- Exporting a model with preprocessing (e.g. normalizing/standardizing image according to normalization parameters during training)
- Exporting a model with postprocessing (e.g. predictions decoding and NMS) - you obtain the ready-to-consume bounding box outputs
- FP16 / INT8 quantization support with calibration
- Pre- and post-processing steps can be customized by the user if needed
- Customising input image shape and batch size
- Customising NMS parameters and number of detections per image
- Customising output format (flat or batched)
Support matrix
It is important to note that different versions of TensorRT has varying support of ONNX opsets.
The support matrix below shows the compatibility of different versions of TensorRT runtime in regard to batch size and output format.
We recommend to use the latest version of TensorRT available.
| Batch Size | Format | OnnxRuntime 1.13.1 | TensorRT 8.4.2 | TensorRT 8.5.3 | TensorRT 8.6.1 |
|------------|--------|--------------------|----------------|----------------|----------------|
| 1 | Flat | Yes | Yes | Yes | Yes |
| >1 | Flat | Yes | Yes | Yes | Yes |
| 1 | Batch | Yes | No | No | Yes |
| >1 | Batch | Yes | No | No | Yes |
python!pip install -qq super-gradients==3.4.0
Minimalistic export example
Let start with the most simple example of exporting a model to ONNX format.
We will use YoloNAS-S model in this example. All models that suports new export API now expose a
export() method that can be used to export a model. There is one mandatory argument that should be passed to the export() method - the path to the output file. Currently, only .onnx format is supported, but we may add support for CoreML and other formats in the future.
pythonfrom super_gradients.common.object_names import Models
from super_gradients.training import modelsmodel = models.get(Models.YOLO_NAS_POSE_S, pretrained_weights="coco_pose")
export_result = model.export("yolo_nas_pose_s.onnx")
A lot of work just happened under the hood:
* A model was exported to ONNX format using default batch size of 1 and input image shape that was used during training
* A preprocessing and postprocessing steps were attached to ONNX graph
* For pre-processing step, the normalization parameters were extracted from the model itself (to be consistent with the image normalization and channel order used during training)
* For post-processing step, the NMS parameters were also extracted from the model and NMS module was attached to the graph
* ONNX graph was checked and simplified to improve compatibility with ONNX runtimes.
A returned value of
export() method is an instance of ModelExportResult class.
First of all it serves the purpose of storing all the information about the exported model in a single place.
It also provides a convenient way to get an example of running the model and getting the output:
pythonexport_result
Model exported successfully to yolo_nas_pose_s.onnx
Model expects input image of shape [1, 3, 640, 640]
Input image dtype is torch.uint8
Exported model already contains preprocessing (normalization) step, so you don't need to do it manually.
Preprocessing steps to be applied to input image are:
Sequential(
(0): CastTensorTo(dtype=torch.float32)
(1): ChannelSelect(channels_indexes=tensor([2, 1, 0]))
(2): ApplyMeanStd(mean=[0.], scale=[255.])
)
Exported model contains postprocessing (NMS) step with the following parameters:
num_pre_nms_predictions=1000
max_predictions_per_image=1000
nms_threshold=0.7
confidence_threshold=0.05
output_predictions_format=batch
Exported model is in ONNX format and can be used with ONNXRuntime
To run inference with ONNXRuntime, please use the following code snippet:
import onnxruntime
import numpy as np
session = onnxruntime.InferenceSession("yolo_nas_pose_s.onnx", providers=["CUDAExecutionProvider", "CPUExecutionProvider"])
inputs = [o.name for o in session.get_inputs()]
outputs = [o.name for o in session.get_outputs()]
example_input_image = np.zeros((1, 3, 640, 640)).astype(np.uint8)
predictions = session.run(outputs, {inputs[0]: example_input_image})
Exported model can also be used with TensorRT
To run inference with TensorRT, please see TensorRT deployment documentation
You can benchmark the model using the following code snippet:
trtexec --onnx=yolo_nas_pose_s.onnx --fp16 --avgRuns=100 --duration=15
Exported model has predictions in batch format:
num_detections, pred_boxes, pred_scores, pred_joints = predictions
for image_index in range(num_detections.shape[0]):
for i in range(num_detections[image_index,0]):
confidence = pred_scores[image_index, i]
x_min, y_min, x_max, y_max = pred_boxes[image_index, i]
pred_joints = pred_joints[image_index, i]
print(f"Detected pose with confidence={confidence}, x_min={x_min}, y_min={y_min}, x_max={x_max}, y_max={y_max}")
for joint_index, (x, y, confidence) in enumerate(pred_joints[i]):
print(f"Joint {joint_index} has coordinates x={x}, y={y}, confidence={confidence}")
That's it. You can now use the exported model with any ONNX-compatible runtime or accelerator.
pythonimport cv2
import numpy as np
from super_gradients.training.utils.media.image import load_image
import onnxruntimeimage = load_image("https://deci-pretrained-models.s3.amazonaws.com/sample_images/beatles-abbeyroad.jpg")
image = cv2.resize(image, (export_result.input_image_shape[1], export_result.input_image_shape[0]))
image_bchw = np.transpose(np.expand_dims(image, 0), (0, 3, 1, 2))
session = onnxruntime.InferenceSession(export_result.output,
providers=["CUDAExecutionProvider", "CPUExecutionProvider"])
inputs = [o.name for o in session.get_inputs()]
outputs = [o.name for o in session.get_outputs()]
result = session.run(outputs, {inputs[0]: image_bchw})
result[0].shape, result[1].shape, result[2].shape, result[3].shape
((1, 1), (1, 1000, 4), (1, 1000), (1, 1000, 17, 3))
In the next section we unpack the result of prediction and show how to use it.
Output format for detection models
If
preprocessing=True (default value) then all models will be exported with NMS. If preprocessing=False models will be exported without NMS and raw model outputs will be returned. In this case, you will need to apply NMS yourself. This is useful if you want to use a custom NMS implementation that is not ONNX-compatible. In most cases you will want to use default preprocessing=True. It is also possible to pass a custom nn.Module as a postprocessing argument to the export() method. This module will be attached to the exported ONNX graph instead of the default NMS module. We encourage users to read the documentation of the export() method to learn more about the advanced options.When exporting an object detection model with postprocessing enabled, the prediction format can be one of two:
* A "flat" format -
DetectionOutputFormatMode.FLAT_FORMAT
* A "batched" format - DetectionOutputFormatMode.BATCH_FORMATYou can select the desired output format by setting
export(..., output_predictions_format=DetectionOutputFormatMode.BATCH_FORMAT).Flat format
A detection results returned as a single tensor of shape
[N, 6 + 3 * NumKeypoints], where N is the number of detected objects in the entire batch. Each row in the tensor represents a single detection result and has the following format:[batch_index, x1, y1, x2, y2, pose confidence, (x,y,score) * num_keypoints]When exporting a model with batch size of 1 (default mode) you can ignore the first column as all boxes will belong to the single sample. In case you export model with batch size > 1 you have to iterate over this array like so:
pythondef iterate_over_flat_predictions(predictions, batch_size):
[flat_predictions] = predictions for image_index in range(batch_size):
mask = flat_predictions[:, 0] == image_index
pred_bboxes = flat_predictions[mask, 1:5]
pred_scores = flat_predictions[mask, 5]
pred_joints = flat_predictions[mask, 6:].reshape((len(pred_bboxes), -1, 3))
yield image_index, pred_bboxes, pred_scores, pred_joints
Iteration over the predictions would be as follows:
pythonfor image_index, pred_bboxes, pred_scores, pred_joints in iterate_over_flat_predictions(predictions, batch_size):
... # Do something useful with the predictions
Batch format
A second supported format is so-called "batch". It matches with output format of TensorRT's NMS implementation. The return value in this case is tuple of 4 tensors:
*
num_predictions - [B, 1] - A number of predictions per sample
* pred_boxes - [B, N, 4] - A coordinates of the predicted boxes in X1, Y1, X2, Y2 format
* pred_scores - [B, N] - A scores of the predicted boxes
* pred_classes - [B, N] - A class indices of the predicted boxesHere
B corresponds to batch size and N is the maximum number of detected objects per image.
In order to get the actual number of detections per image you need to iterate over num_predictions tensor and get the first element of each row.A corresponding code snippet for iterating over the batch predictions would look like this:
pythondef iterate_over_batch_predictions(predictions, batch_size):
num_detections, batch_boxes, batch_scores, batch_joints = predictions
for image_index in range(batch_size):
num_detection_in_image = num_detections[image_index, 0] pred_scores = batch_scores[image_index, :num_detection_in_image]
pred_boxes = batch_boxes[image_index, :num_detection_in_image]
pred_joints = batch_joints[image_index, :num_detection_in_image].reshape((len(pred_scores), -1, 3))
yield image_index, pred_boxes, pred_scores, pred_joints
And similary to flat format, iteration over the predictions would be as follows:
pythonfor image_index, pred_bboxes, pred_scores, pred_joints in iterate_over_batch_predictions(predictions, batch_size):
... # Do something useful with the predictions
Now when you're familiar with the output formats, let's see how to use them.
To start, it's useful to take a look at the values of the predictions with a naked eye:
pythonnum_predictions, pred_boxes, pred_scores, pred_poses = result
num_predictions
array([[9]], dtype=int64)
pythonnp.set_printoptions(threshold=3, edgeitems=3)
pred_boxes, pred_boxes.shape
(array([[[182.49644 , 249.07802 , 305.27576 , 530.3644 ],
[ 34.52883 , 247.74242 , 175.7783 , 544.1926 ],
[438.808 , 251.08049 , 587.11865 , 552.69336 ],
...,
[ 67.20265 , 248.3974 , 122.415375, 371.65637 ],
[625.7083 , 306.74194 , 639.4926 , 501.08337 ],
[450.61108 , 386.74622 , 556.77325 , 523.2412 ]]], dtype=float32),
(1, 1000, 4))
pythonnp.set_printoptions(threshold=3, edgeitems=3)
pred_scores, pred_scores.shape
(array([[0.84752125, 0.826281 , 0.82436883, ..., 0.00848398, 0.00848269,
0.00848123]], dtype=float32),
(1, 1000))
pythonnp.set_printoptions(threshold=3, edgeitems=3)
pred_poses, pred_poses.shape
(array([[[[2.62617737e+02, 2.75986389e+02, 7.74692297e-01],
[2.63401123e+02, 2.70397522e+02, 3.57395113e-01],
[2.57980499e+02, 2.70888336e+02, 7.75521040e-01],
...,
[2.58518188e+02, 4.50223969e+02, 9.40084636e-01],
[2.01152466e+02, 5.02089630e+02, 8.42420936e-01],
[2.82095978e+02, 5.06688324e+02, 8.73963714e-01]],
[[1.14750252e+02, 2.75872864e+02, 8.29551518e-01],
[1.15829544e+02, 2.70712891e+02, 4.48927283e-01],
[1.09389343e+02, 2.70643494e+02, 8.33203077e-01],
...,
[7.29626541e+01, 4.55435028e+02, 9.07496691e-01],
[1.47440369e+02, 5.05209564e+02, 8.53177905e-01],
[5.24395561e+01, 5.16123291e+02, 8.44702840e-01]],
[[5.46199341e+02, 2.83605713e+02, 6.09813333e-01],
[5.45253479e+02, 2.78786011e+02, 1.59033239e-01],
[5.44112183e+02, 2.78675476e+02, 5.77503145e-01],
...,
[5.00366119e+02, 4.57584869e+02, 8.84028912e-01],
[5.50320129e+02, 5.21863281e+02, 7.15586364e-01],
[4.54590271e+02, 5.17590332e+02, 7.93488443e-01]],
...,
[[1.13875908e+02, 2.76212708e+02, 7.35527277e-01],
[1.16164986e+02, 2.70696411e+02, 4.00955290e-01],
[1.08107491e+02, 2.70656555e+02, 7.91907310e-01],
...,
[9.75953293e+01, 4.07489868e+02, 3.45197320e-01],
[1.01579475e+02, 4.40818176e+02, 2.17337132e-01],
[9.04172211e+01, 4.44152771e+02, 2.28111655e-01]],
[[6.42500244e+02, 3.39081055e+02, 1.75797671e-01],
[6.42386841e+02, 3.34906342e+02, 1.55016124e-01],
[6.41675354e+02, 3.34820374e+02, 1.29657656e-01],
...,
[6.40000122e+02, 4.15383392e+02, 2.22081602e-01],
[6.37456421e+02, 4.40941406e+02, 2.00318485e-01],
[6.39243164e+02, 4.41459686e+02, 2.33620048e-01]],
[[5.17478271e+02, 4.09209961e+02, 1.95783913e-01],
[5.21710632e+02, 4.01950928e+02, 1.90346301e-01],
[5.12909302e+02, 4.02274841e+02, 1.88751698e-01],
...,
[4.98697205e+02, 4.55512695e+02, 4.54110742e-01],
[5.19384705e+02, 5.21536316e+02, 4.20579553e-01],
[4.83649933e+02, 5.19510498e+02, 4.25356269e-01]]]],
dtype=float32),
(1, 1000, 17, 3))
Visualizing predictions
For sake of this tutorial we will use a simple visualization function that is tailored for batch_size=1 only.
You can use it as a starting point for your own visualization code.
pythonfrom super_gradients.training.utils.visualization.pose_estimation import PoseVisualization
import matplotlib.pyplot as pltdef show_predictions_from_batch_format(image, predictions):
# In this tutorial we are using batch size of 1, therefore we are getting only first element of the predictions
image_index, pred_boxes, pred_scores, pred_joints = next(iter(iterate_over_batch_predictions(predictions, 1)))
image = PoseVisualization.draw_poses(
image=image, poses=pred_joints, scores=pred_scores, boxes=pred_boxes,
edge_links=None, edge_colors=None, keypoint_colors=None, is_crowd=None
)
plt.figure(figsize=(8, 8))
plt.imshow(image)
plt.tight_layout()
plt.show()
pythonshow_predictions_from_batch_format(image, result)
Changing the output format
You can explicitly specify output format of the predictions by setting the
output_predictions_format argument of export() method. Let's see how it works:pythonfrom super_gradients.conversion import DetectionOutputFormatModeexport_result = model.export("yolo_nas_s.onnx", output_predictions_format=DetectionOutputFormatMode.FLAT_FORMAT)
export_result
Model exported successfully to yolo_nas_s.onnx
Model expects input image of shape [1, 3, 640, 640]
Input image dtype is torch.uint8
Exported model already contains preprocessing (normalization) step, so you don't need to do it manually.
Preprocessing steps to be applied to input image are:
Sequential(
(0): CastTensorTo(dtype=torch.float32)
(1): ChannelSelect(channels_indexes=tensor([2, 1, 0]))
(2): ApplyMeanStd(mean=[0.], scale=[255.])
)
Exported model contains postprocessing (NMS) step with the following parameters:
num_pre_nms_predictions=1000
max_predictions_per_image=1000
nms_threshold=0.7
confidence_threshold=0.05
output_predictions_format=flat
Exported model is in ONNX format and can be used with ONNXRuntime
To run inference with ONNXRuntime, please use the following code snippet:
import onnxruntime
import numpy as np
session = onnxruntime.InferenceSession("yolo_nas_s.onnx", providers=["CUDAExecutionProvider", "CPUExecutionProvider"])
inputs = [o.name for o in session.get_inputs()]
outputs = [o.name for o in session.get_outputs()]
example_input_image = np.zeros((1, 3, 640, 640)).astype(np.uint8)
predictions = session.run(outputs, {inputs[0]: example_input_image})
Exported model can also be used with TensorRT
To run inference with TensorRT, please see TensorRT deployment documentation
You can benchmark the model using the following code snippet:
trtexec --onnx=yolo_nas_s.onnx --fp16 --avgRuns=100 --duration=15
Exported model has predictions in flat format:
# flat_predictions is a 2D array of [N,K] shape
# Each row represents (image_index, x_min, y_min, x_max, y_max, confidence, joints...)
# Please note all values are floats, so you have to convert them to integers if needed
[flat_predictions] = predictions
pred_bboxes = flat_predictions[:, 1:5]
pred_scores = flat_predictions[:, 5]
pred_joints = flat_predictions[:, 6:].reshape((len(pred_bboxes), -1, 3))
for i in range(len(pred_bboxes)):
confidence = pred_scores[i]
x_min, y_min, x_max, y_max = pred_bboxes[i]
print(f"Detected pose with confidence={{confidence}}, x_min={{x_min}}, y_min={{y_min}}, x_max={{x_max}}, y_max={{y_max}}")
for joint_index, (x, y, confidence) in enumerate(pred_joints[i]):")
print(f"Joint {{joint_index}} has coordinates x={{x}}, y={{y}}, confidence={{confidence}}")
Now we exported a model that produces predictions in
flat format. Let's run the model like before and see the result:
pythonsession = onnxruntime.InferenceSession(export_result.output,
providers=["CUDAExecutionProvider", "CPUExecutionProvider"])
inputs = [o.name for o in session.get_inputs()]
outputs = [o.name for o in session.get_outputs()]
result = session.run(outputs, {inputs[0]: image_bchw})
result[0].shape
(9, 57)
pythondef show_predictions_from_flat_format(image, predictions):
image_index, pred_boxes, pred_scores, pred_joints = next(iter(iterate_over_flat_predictions(predictions, 1))) image = PoseVisualization.draw_poses(
image=image, poses=pred_joints, scores=pred_scores, boxes=pred_boxes,
edge_links=None, edge_colors=None, keypoint_colors=None, is_crowd=None
)
plt.figure(figsize=(8, 8))
plt.imshow(image)
plt.tight_layout()
plt.show()
pythonshow_predictions_from_flat_format(image, result)
Changing postprocessing settings
You can control a number of parameters in the NMS settings as well as maximum number of detections per image before and after NMS step:
* IOU threshold for NMS -
nms_iou_threshold
* Score threshold for NMS - nms_score_threshold
* Maximum number of detections per image before NMS - max_detections_before_nms
* Maximum number of detections per image after NMS - max_detections_after_nmsFor sake of demonstration, let's export a model that would produce at most one detection per image with confidence threshold above 0.8 and NMS IOU threshold of 0.5. Let's use at most 100 predictions per image before NMS step:
pythonexport_result = model.export(
"yolo_nas_s_pose_top_1.onnx",
confidence_threshold=0.8,
nms_threshold=0.5,
num_pre_nms_predictions=100,
max_predictions_per_image=1,
output_predictions_format=DetectionOutputFormatMode.FLAT_FORMAT
)session = onnxruntime.InferenceSession(export_result.output,
providers=["CUDAExecutionProvider", "CPUExecutionProvider"])
inputs = [o.name for o in session.get_inputs()]
outputs = [o.name for o in session.get_outputs()]
result = session.run(outputs, {inputs[0]: image_bchw})
show_predictions_from_flat_format(image, result)
As expected, the predictions contains exactly one detection with the highest confidence score.
Export of quantized model
You can export a model with quantization to FP16 or INT8. To do so, you need to specify the
quantization_mode argument of export() method.Important notes:
* Quantization to FP16 requires CUDA / MPS device available and would not work on CPU-only machines.
Let's see how it works:
pythonfrom super_gradients.conversion.conversion_enums import ExportQuantizationModeexport_result = model.export(
"yolo_nas_pose_s_int8.onnx",
confidence_threshold=0.5,
output_predictions_format=DetectionOutputFormatMode.FLAT_FORMAT,
quantization_mode=ExportQuantizationMode.INT8 # or ExportQuantizationMode.FP16
)
session = onnxruntime.InferenceSession(export_result.output,
providers=["CUDAExecutionProvider", "CPUExecutionProvider"])
inputs = [o.name for o in session.get_inputs()]
outputs = [o.name for o in session.get_outputs()]
result = session.run(outputs, {inputs[0]: image_bchw})
show_predictions_from_flat_format(image, result)
Advanced INT-8 quantization options
When quantizing a model using
quantization_mode==ExportQuantizationMode.INT8 you can pass a DataLoader to export() function to collect correct statistics of activations to prodice a more accurate quantized model.
We expect the DataLoader to return either a tuple of tensors or a single tensor. In case a tuple of tensors is returned by data-loader the first element will be used as input image.
You can use existing data-loaders from SG here as is.Important notes
* A
calibration_loader should use same image normalization parameters that were used during training.In the example below we use a dummy data-loader for sake of showing how to use this feature. You should use your own data-loader here.
pythonimport torch
from torch.utils.data import DataLoader
from super_gradients.conversion import ExportQuantizationModeTHIS IS ONLY AN EXAMPLE. YOU SHOULD USE YOUR OWN DATA-LOADER HERE
dummy_calibration_dataset = [torch.randn((3, 640, 640), dtype=torch.float32) for _ in range(32)]
dummy_calibration_loader = DataLoader(dummy_calibration_dataset, batch_size=8, num_workers=0)
THIS IS ONLY AN EXAMPLE. YOU SHOULD USE YOUR OWN DATA-LOADER HERE
export_result = model.export(
"yolo_nas_pose_s_int8_with_calibration.onnx",
confidence_threshold=0.5,
output_predictions_format=DetectionOutputFormatMode.FLAT_FORMAT,
quantization_mode=ExportQuantizationMode.INT8,
calibration_loader=dummy_calibration_loader
)
session = onnxruntime.InferenceSession(export_result.output,
providers=["CUDAExecutionProvider", "CPUExecutionProvider"])
inputs = [o.name for o in session.get_inputs()]
outputs = [o.name for o in session.get_outputs()]
result = session.run(outputs, {inputs[0]: image_bchw})
show_predictions_from_flat_format(image, result)
25%|βββββββββββββββββββββββββββββββββ | 4/16 [00:12<00:37, 3.10s/it]
Limitations
* Dynamic batch size / input image shape is not supported yet. You can only export a model with a fixed batch size and input image shape.
* TensorRT of version 8.5.2 or higher is required.
* Quantization to FP16 requires CUDA / MPS device available.
Conclusion
This concludes the export tutorial for YoloNAS-Pose pose estimation model.
We hope you found it useful and will be able to use it to export your own models to ONNX format.
In case you have any questions or issues, please feel free to reach out to us at https://github.com/Deci-AI/super-gradients/issues.
---
Documentation/Source/ObjectDetection (documentation/source/ObjectDetection.md)
Object Detection
Object detection is a core task in computer vision that allows to detect and classify bounding boxes in images.
It's been gaining popularity and ubiquity extremely fast since the first breakthroughs in Deep Learning and advanced a wide range of companies, including the medical domain, surveillance, smart shopping, etc.
It comes as no surprise considering that it covers two basic needs in an end-to-end manner: to find all present objects and to assign a class to each one of them,
while cleverly dealing with the background and its dominance over all other classes.
Due to this, most recent research publications dedicated to object detection focus on a good trade-off between accuracy and speed.
In SuperGradients, we aim to collect such models and make them very convenient and accessible to you, so that you can try any one of them interchangeably.
Implemented models
| Model | Yaml | Model class | Loss Class | NMS Callback |
|----------------------------------------------|--------------------------------------------------------------------------------------------------------------------------------------------------------------------------|------------------------------------------------------------------------------------------------------------------------------------------|---------------------------------------------------------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| SSD | ssd_lite_mobilenetv2_arch_params | SSDLiteMobileNetV2 | SSDLoss | SSDPostPredictCallback |
| YOLOX | yolox_s_arch_params | YoloX_S | YoloXFastDetectionLoss | YoloXPostPredictionCallback |
| PPYolo | ppyoloe_arch_params | PPYoloE | PPYoloELoss | PPYoloEPostPredictionCallback |
| YoloNAS | yolo_nas_s_arch_params | Yolo NAS S | PPYoloELoss | PPYoloEPostPredictionCallback |
Datasets
There are several well-known datasets for object detection: COCO, Pascal, etc.
SuperGradients provides ready-to-use dataloaders for the COCO dataset COCODetectionDataset
and more general
DetectionDataset implementation that you can subclass from for your specific dataset format.If you want to load the dataset outside of a yaml training, do:
pythonfrom super_gradients.training import dataloaders
data_dir = "/path/to/coco_dataset_dir"
train_dataloader = dataloaders.get(name='coco2017_train',
dataset_params={"data_dir": data_dir},
dataloader_params={'num_workers': 2}
)
val_dataloader = dataloaders.get(name='coco2017_val',
dataset_params={"data_dir": data_dir},
dataloader_params={'num_workers': 2}
)
Loss functions
Generally speaking, in object detection task the loss function is tightly coupled with the model and cannot be used interchangeably. E.g. you cannot use YoloX loss with YoloNAS model and vice versa.
This is different from classification or segmentation task where model output is "standard" and usually does not change.
In Object Detection task, the model output format may vary greatly and also training objective is often tailored for a specific model architecture.
To indicate compatibility between a model and a loss function, we use the convention of model name and loss function starting from the same prefix name. For example:
SSDLiteMobileNetV2 model & SSDLoss, YoloX_S and YoloXFastDetectionLoss, etc.Of course, you are free to adjust hyperparameters of the loss function to your liking. Let's check a
PPYoloELoss loss class as an example:
It has the following constructor:python@register_loss(Losses.PPYOLOE_LOSS)
class PPYoloELoss(nn.Module):
def __init__(
self,
num_classes: int,
use_varifocal_loss: bool = True,
use_static_assigner: bool = True,
reg_max: int = 16,
classification_loss_weight: float = 1.0,
iou_loss_weight: float = 2.5,
dfl_loss_weight: float = 0.5,
):
...
In your recipe you can pass the desired values for each parameter.
For example show below, we increase the classification component weight to 10 and set the DFL & IOU components of the loss to 1.0:
yamltraining_hyperparams:
loss:
ppyoloe_loss:
num_classes: ${arch_params.num_classes}
classification_loss_weight: 10
iou_loss_weight: 1.0
dfl_loss_weight: 1.0
This is how you can modify the loss hyperparameters. If you need to modify the loss itself, you can subclass it and override the
forward method to fit your needs.python@register_loss()
class MyCustomPPYoloELoss(nn.Module):
def forward(self, outputs, target):
...
yamltraining_hyperparams:
loss: MyCustomPPYoloELoss
criterion_params:
num_classes: ${arch_params.num_classes}
classification_loss_weight: 10
iou_loss_weight: 1.0
dfl_loss_weight: 1.0
Metrics
A typical metric for object detection is mean average precision, mAP for short.
It is calculated for a specific IoU level which defines how tightly a predicted box must intersect with a ground truth box to be considered a true positive.
Both one value and a range can be used as IoU, where a range refers to an average of mAPs for each IoU level.
The most popular metric for mAP on COCO is [email protected]:0.95, SuperGradients provides its implementation DetectionMetrics.
It is written to be as close as possible to the official metric implementation from COCO API, while being much faster and DDP-friendly.
We provide a few metrics for object detection with pre-defined IoU levels to fit the most frequent use cases:
* DetectionMetrics_050_095 - computes mAP at IoU range [0.5; 0.95] with a step of 0.05 (Default COCO metric)
* DetectionMetrics_050 - computes mAP at IoU level 0.5
* DetectionMetrics_075 - computes mAP at IoU level 0.75
* DetectionMetrics - computes mAP at user-specified IoU level (Defaults to [0.5; 0.95])
You can also specify a custom IoU range or a single IoU level for the metric.
In addition to computing mAP,
DetectionMetrics also computes other metrics such as:* Recall score at a given score threshold
* Precision score at a given score threshold
* F-1 detection score at a given score threshold
* Average precision score for each class
DetectionMetrics can even find the optimal confidence threshold that maximizes mean F1 score.
Here is how to enable computing all these metrics:
yamltraining_hyperparams:
valid_metrics_list:
- DetectionMetrics:
num_cls: ${num_classes}
normalize_targets: True
score_thres: 0.1 # A lower bound rejection threshold for predictions
top_k_predictions: 300 # At most 300 predictions per image will be considered with confidence above score_thres
iou_thres: [0.6, 0.8] # <--- IoU range [0.6; 0.8] with 0.05 step
include_classwise_ap: True # Enables computing AP for each class (helps to find problematic classes)
calc_best_score_thresholds: True # Enables computing optimal confidence threshold that maximizes mean F1 score
post_prediction_callback:
_target_: super_gradients.training.models.detection_models.pp_yolo_e.PPYoloEPostPredictionCallback
score_threshold: 0.01
nms_top_k: 1000
max_predictions: 300
nms_threshold: 0.7 metric_to_watch: '[email protected]:0.80'
In order to use
DetectionMetrics you have to pass a so-called post_prediction_callback to the metric, which is responsible for the postprocessing of the model's raw output into final predictions and is explained below. Postprocessing
Postprocessing refers to a process of transforming the model's raw output into final predictions. Postprocessing is also model-specific and depends on the model's output format.
For
YOLOX model, the postprocessing step is implemented in YoloXPostPredictionCallback class.
It can be passed into a DetectionMetrics as a post_prediction_callback.
The postprocessing of all detection models involves non-maximum suppression (NMS) which filters dense model's predictions and leaves only boxes with the highest confidence and suppresses boxes with very high overlap
based on the assumption that they likely belong to the same object. Thus, a confidence threshold and an IoU threshold must be passed into the postprocessing object.pythonfrom super_gradients.training.models.detection_models.yolo_base import YoloXPostPredictionCallback
post_prediction_callback = YoloXPostPredictionCallback(conf=0.001, iou=0.6)
All post prediction callbacks returns a list of lists with decoded boxes after NMS:
List[torch.Tensor].
The first list wraps all images in the batch, and each tensor holds all predictions for each image in the batch.
The shape of predictions tensor is [N, 6] where N is the number of predictions for the image and each row is holds values of [X1, Y1, X2, Y2, confidence, class_id].Box coordinates are in absolute (pixel) units.
Visualization
Visualization of the model predictions is a very important part of the training process for any computer vision task.
By visualizing the predicted boxes, developers and researchers can identify errors or inaccuracies in the model's output and adjust the model's architecture or training data accordingly.
#### Extreme Batch Visualization during training
SuperGradients provides an implementation of ExtremeBatchDetectionVisualizationCallback.
You can use this callback in your training pipeline to visualize best or worst batch during training.
This callback observes a specific metric during training epoch and logs the most extreme batch to configured logger (Default is Tensorboard).
The logging includes visualization of ground truth boxes and model's predictions.
To use this callback you would need to add it to
training_hyperparams.phase_callbacks in your yaml:yamltraining_hyperparams:
phase_callbacks:
- ExtremeBatchDetectionVisualizationCallback:
metric: # Defines which metric to observe
DetectionMetrics_050:
score_thres: 0.1
top_k_predictions: 300
num_cls: ${num_classes}
normalize_targets: True
post_prediction_callback:
_target_: super_gradients.training.models.detection_models.pp_yolo_e.PPYoloEPostPredictionCallback
score_threshold: 0.01
nms_top_k: 1000
max_predictions: 300
nms_threshold: 0.7
max: False # Indicates that we want to log batch with the lowest metric value
metric_component_name: '[email protected]'
post_prediction_callback:
_target_: super_gradients.training.models.detection_models.pp_yolo_e.PPYoloEPostPredictionCallback
score_threshold: 0.25
nms_top_k: 1000
max_predictions: 300
nms_threshold: 0.7
normalize_targets: True
Note in the example below the ExtremeBatchDetectionVisualizationCallback callback observes a user-provided Metric class that computes the score for each batch.
You may also observe the entire loss or individual components of the loss as follows.
In this case instead of passing metric argument to constructor of ExtremeBatchDetectionVisualizationCallback you would need to pass loss_to_monitor argument.
The fully qualified name of the loss includes the loss class name and component name separated by /:yamltraining_hyperparams:
phase_callbacks:
- ExtremeBatchDetectionVisualizationCallback:
loss_to_monitor: "YoloNASPoseLoss/loss"
max: True
#### Visualization of predictions after training
pythonimport torch
import numpy as npfrom super_gradients.training import models
from super_gradients.training.utils.detection_utils import DetectionVisualization
from super_gradients.training.datasets.datasets_conf import COCO_DETECTION_CLASSES_LIST
def my_undo_image_preprocessing(im_tensor: torch.Tensor) -> np.ndarray:
im_np = im_tensor.cpu().numpy()
im_np = im_np[:, ::-1, :, :].transpose(0, 2, 3, 1)
im_np *= 255.0
return np.ascontiguousarray(im_np, dtype=np.uint8)
model = models.get("yolox_s", pretrained_weights="coco", num_classes=80)
imgs, targets = next(iter(train_dataloader))
preds = model.get_post_prediction_callback(conf=0.1, iou=0.6)(model(imgs))
DetectionVisualization.visualize_batch(imgs, preds, targets, batch_name='train', class_names=COCO_DETECTION_CLASSES_LIST,
checkpoint_dir='/path/for/saved_images/', gt_alpha=0.5,
undo_preprocessing_func=my_undo_image_preprocessing)
The function you pass as undo_preprocessing_func will define how to undo dataset transforms and return the image back into its initial formal (BGR, uint8).
This also allows you to test the correctness of your dataset implementation, since it saves images after they go through transforms. This may be especially useful for a train set with heavy augmentation transforms. You can see both the predictions and the ground truth, and give the ground truth box the desired opacity. The saved train image for a dataset with a mosaic transform should look something like this:
#### Visualization of predictions after training using predict()
If you would like to do the visualization outside of training you can use
predict() method that is implemented for most of our detection models.
pythonmodel = models.get("yolox_s", pretrained_weights="coco", num_classes=80)
model.predict("https://deci-pretrained-models.s3.amazonaws.com/sample_images/beatles-abbeyroad.jpg").show()
See for more details on using Predict API.
Let's train!
As stated above, training can be launched with just one command. For the curious ones, let's see how all the components we've just discussed fall into place in one yaml.
yamlcoco2017_yolox
defaults:
- training_hyperparams: coco2017_yolox_train_params
- dataset_params: coco_detection_dataset_params
- arch_params: yolox_s_arch_params
- checkpoint_params: default_checkpoint_params
- _self_
These are the actual components of coco2017_yolox.yaml
The dataset parameters are defined in dataset_params: and are eventually passed into coco2017_train/val dataset mentioned above in the Datasets section The metric is part of
training_hyperparams and so it's stated in the coco2017_yolox_train_params.yaml with:yamlvalid_metrics_list:
- DetectionMetrics:
normalize_targets: True
post_prediction_callback:
_target_: super_gradients.training.models.detection_models.yolo_base.YoloXPostPredictionCallback
iou: 0.65
conf: 0.01
num_cls: 80
Notice how
YoloXPostPredictionCallback is passed as a post_prediction_callback.A visualization belongs to
training_hyperparams as well, specifically to the phase_callbacks list, as follows:
yamlphase_callbacks:
- DetectionVisualizationCallback:
phase:
_target_: super_gradients.training.utils.callbacks.callbacks.Phase
value: VALIDATION_EPOCH_END
freq: 1
post_prediction_callback:
_target_: super_gradients.training.models.detection_models.yolo_base.YoloXPostPredictionCallback
iou: 0.65
conf: 0.01
classes: [
"person", "bicycle", "car", "motorcycle", "airplane", "bus", "train", "truck", "boat", "traffic light",
"fire hydrant", "stop sign", "parking meter", "bench", "bird", "cat", "dog", "horse", "sheep", "cow",
"elephant", "bear", "zebra", "giraffe", "backpack", "umbrella", "handbag", "tie", "suitcase", "frisbee",
"skis", "snowboard", "sports ball", "kite", "baseball bat", "baseball glove", "skateboard", "surfboard",
"tennis racket", "bottle", "wine glass", "cup", "fork", "knife", "spoon", "bowl", "banana", "apple",
"sandwich", "orange", "broccoli", "carrot", "hot dog", "pizza", "donut", "cake", "chair", "couch",
"potted plant", "bed", "dining table", "toilet", "tv", "laptop", "mouse", "remote", "keyboard", "cell phone",
"microwave", "oven", "toaster", "sink", "refrigerator", "book", "clock", "vase", "scissors", "teddy bear",
"hair drier", "toothbrush"
]
By default, this callback is not part of the yaml, so you can add it yourself if you prefer.Using the provided yaml, SuperGradients can instantiate all the components and launch training from config:
pythontrainer = Trainer(experiment_name=cfg.experiment_name, ckpt_root_dir=cfg.ckpt_root_dir)BUILD NETWORK
model = models.get(
model_name=cfg.architecture,
num_classes=cfg.arch_params.num_classes,
arch_params=cfg.arch_params,
strict_load=cfg.checkpoint_params.strict_load,
pretrained_weights=cfg.checkpoint_params.pretrained_weights,
checkpoint_path=cfg.checkpoint_params.checkpoint_path,
load_backbone=cfg.checkpoint_params.load_backbone,
)INSTANTIATE DATA LOADERS
train_dataloader = dataloaders.get(
name=get_param(cfg, "train_dataloader"),
dataset_params=cfg.dataset_params.train_dataset_params,
dataloader_params=cfg.dataset_params.train_dataloader_params,
)
val_dataloader = dataloaders.get(
name=get_param(cfg, "val_dataloader"),
dataset_params=cfg.dataset_params.val_dataset_params,
dataloader_params=cfg.dataset_params.val_dataloader_params,
)
recipe_logged_cfg = {"recipe_config": OmegaConf.to_container(cfg, resolve=True)}
TRAIN
res = trainer.train(
model=model,
train_loader=train_dataloader,
valid_loader=val_dataloader,
training_params=cfg.training_hyperparams,
additional_configs_to_log=recipe_logged_cfg,
)
It is convenient to trigger it with train_from_recipe.py, but you can do the same in your project by constructing the desired objects directly.
How to connect your own dataset
To add a new dataset to SuperGradients, you need to implement a few things:
- Implement a new dataset class
- Add a configuration file
Let's unwrap each of the steps
Implement a new dataset class
To train an existing architecture on a new dataset one needs to implement the dataset class first.
It is generally a good idea to subclass from
DetectionDataset as it comes with a few useful features,
such as subclassing, caching, extra sample loading necessary for complex transform like mosaic or mixup, etc. It requires you to implement only a few methods for files loading.
If you prefer, you can use torch.utils.data.Dataset as well.A minimal implementation of a
DetectionDataset subclass class should look similar to this:text/ Detailed source-code truncated for AI context efficiency. /
Note the addition of @register_dataset decorator. This makes SuperGradients recognize your dataset so that you can use its name directly in a yaml.
Since detection labels often contain different number of boxes per image, targets are padded with 0s, which allows to use them in a batch.
They are later removed by a DetectionCollateFN which prepends all targets with an index in a batch and stacks them together.
Add a configuration file
Create new
my_new_dataset_params.yaml file under dataset_params folder and add your dataset and dataloader parameters:text/ Detailed source-code truncated for AI context efficiency. /
In your training recipe add/change the following lines to:
yamlmy_train_recipe.yaml
defaults:
- training_hyperparams: ...
- dataset_params: my_new_dataset_params
- arch_params: ...
- checkpoint_params: ...
- _self_
train_dataloader:
val_dataloader:
num_classes: 3
...
And you should be good to go!
Understanding model's predictions
This section covers what is the output of each model class in train, eval and tracing modes. A tracing mode is enabled
when exporting model to ONNX or when using
torch.jit.trace() call
Corresponding loss functions and post-prediction callbacks from the table above are written to match the output format of the models.
That being said, if you're using YoloX model, you should use YoloX loss and post-prediction callback for YoloX model.
Mixing them with other models will result in an error.It is important to understand the output of the model class in order to use it correctly in the training process and especially
if you are going to use the model's prediction in a custom callback or loss.
YoloX
#### Training modeIn training mode, YoloX returns a list of 3 tensors that contains the intermediates required for the loss calculation.
They correspond to output feature maps of the prediction heads:
- Output feature map at index 0:
[B, 1, H/8, W/8, C + 5]
- Output feature map at index 1: [B, 1, H/16, W/16, C + 5]
- Output feature map at index 2: [B, 1, H/32, W/32, C + 5]Value
C corresponds to the number of classes in the dataset.
And remaining 5elements are box coordinates and objectness score.
Layout of elements in the last dimension is as follows: [cx, cy, w, h, obj_score, class_scores...]
Box regression in these outputs are NOT in pixel coordinates.
X and Y coordinates are normalized coordinates.
Width and height values are the power factor for the base of eoutput_feature_map_at_index_0, output_feature_map_at_index_1, output_feature_map_at_index_2 = yolo_x_model(images)In this mode, predictions decoding is not performed.
#### Eval mode
In eval mode, YoloX returns a tuple of decoded predictions and raw intermediates.
predictions, (raw_predictions_0, raw_predictions_1, raw_predictions_2) = yolo_x_model(images)predictions is a single tensor of shape [B, num_predictions, C + 5] where num_predictions is the total number of predictions across all 3 output feature maps.The layout of the last dimension is the same as in training mode:
[cx, cy, w, h, obj_score, class_scores...].
Values of cx, cy, w, h are in absolute pixel coordinates and confidence scores are in range [0, 1].#### Tracing mode
Same as in Eval mode.
PPYolo-E & Yolo-NAS
#### Training & Validation modePPYoloE & Yolo-NAS returns a tuple of 2 tensors:
decoded_predictions, raw_intermediates. A
decoded_predictions itself is a tuple of 2 tensors ([B,Anchors,4] and [B,Anchors,C]) with decoded bounding boxes and class scores.
A raw_intermediates contains 6 tensors of intermediates required for the loss calculation.You can access individual components of the model's output using the following snippet:
(pred_bboxes, pred_scores), (cls_score_list, reg_distri_list, anchors, anchor_points, num_anchors_list, stride_tensor) = model(images)Here
pred_bboxes and pred_scores are decoded predictions of the model: *
pred_bboxes - [B, num_anchors, 4] - decoded bounding boxes in the format [x1, y1, x2, y2] in absolute (pixel) coordinates
* pred_scores - [B, num_anchors, num_classes] - class scores (0..1) for each bounding boxPlease note that box predictions are not clipped and may extend beyond the image boundaries.
Additionally, the NMS is not performed yet at this stage. This is where the post-prediction callback comes into play.
Remaining tensors contains the intermediates required for the loss calculation. They are as follows:
*
cls_score_list - [B, num_anchors, num_classes]
* reg_distri_list - [B, num_anchors, num_regression_dims]
* anchors - [num_anchors, 4]
* anchor_points - [num_anchors, 2]
* num_anchors_list - [num_anchors]
* stride_tensor - [num_anchors]
#### Tracing mode
In tracing mode, PPYoloE returns only decoded predictions:
pred_bboxes, pred_scores = yolo_nas_model(images)Please note that box predictions are not clipped and may extend beyond the image boundaries.
Additionally, the NMS is not performed yet at this stage. This is where the post-prediction callback comes into play.
Training
The easiest way to start training any mode in SuperGradients is to use a pre-defined recipe. In this tutorial, we will see how to train
YOLOX-S model, other models can be trained by analogy.Prerequisites
1. You have to install SuperGradients first. Please refer to the Installation section for more details.
2. Prepare the COCO dataset as described in the Computer Vision Datasets Setup under Detection Datasets section.
After you meet the prerequisites, you can start training the model by running from the root of the repository:
Training from recipe
bashpython -m super_gradients.train_from_recipe --config-name=coco2017_yolox multi_gpu=Off num_gpus=1
Note, the default configuration for this recipe is to use 8 GPUs in DDP mode. This hardware configuration may not be for everyone, so in the example above we override GPU settings to use a single GPU.
It is highly recommended to read through the recipe file coco2017_yolox to get better understanding of the hyperparameters we use here.
If you're unfamiliar with config files, we recommend you to read the Configuration Files part first.
How to add a new model
To implement a new model, you need to add the following parts:
- Model architecture itself
- Postprocessing Callback
For a custom model, a good starting point would be a
CustomizableDetector
class since it allows to configure a backbone, a neck and a head separately. See an example yaml of
a model that uses it: ssd_lite_mobilenetv2_arch_params
It is strongly advised to use the existing callbacks and to define your model's head such that it returns the same outputs.
---
Documentation/Source/Optimizers (documentation/source/optimizers.md)
Optimizers
Optimization is a critical step in the deep learning process as it determines how well the network will learn from the training data.
SuperGradients supports out-of-the-box pytorch optimizers(
SGD,
Adam,
AdamW,
RMS_PROP), but also
RMSpropTF and
Lamb.
Set the optimizer in the code
Optimizers should be part of the training parameters.
pyfrom super_gradients import Trainertrainer = Trainer(...)
trainer.train(
training_params={"optimizer": "Adam", "optimizer_params": {"eps": 1e-3}, ...},
...
)
Note:
The
optimizer_params is a dictionary of all the optimizer parameters you want to set. It can be any argument defined in the optimizer __init__ method , except for params because this argument corresponds to the model to optimize and is automatically provided by the Trainer.
Set the optimizer in the recipes
When working with recipes, you need to modify the recipes/training_hyperparams as below:yamlrecipes/training_hyperparams/my_training_recipe.yaml
...
optimizer: Adam
optimizer_params:
eps: 1e-3
Use Custom Optimizers
If your own optimizer is not natively supported by SuperGradients, you can always register it!pyfrom super_gradients.common.registry.registry import register_optimizer@register_optimizer()
class CustomOptimizer:
def __init__(
self,
params, # This arg is the only required regardless of your optimizer, the rest depends on your optimizer.
alpha: float,
betas: float
):
defaults = dict(alpha=alpha, betas=betas)
super(CustomOptimizer, self).__init__(params, defaults)
...
And then update your training hyperparameters:
yamlmy_training_hyperparams.yaml
...
optimizer: CustomOptimizer
optimizer_params:
alpha: 1e-3
betas: 1e-3
Customize learning rate for different model blocks
You can define the learning rate to use on each section of your model by working with initialize_param_groups and update_param_groups.
- initialize_param_groups defines the groups, and the learning rate to use for each group. It is called on instantiation.
- update_param_groups updates the learning rate for each group. It is called by LR callbacks (such as LRCallbackBase) during the training. If your model (i.e. any
torch.nn.Module) is lacking these methods, the same learning rate will be applied to every block.
But if you implement them, it will be taken into account by the Trainer just like with any other SuperGradients model.#### Example
Assuming that you have your own custom model and that you want work with a different learning rate on the backbone.
You first need to implement the
initialize_param_groups and update_param_groups accordingly.
pyimport torch
from super_gradients.common.registry.registry import register_model
@register_model() # Required if working with recipe
class MyModel(torch.nn.Module):
...
def initialize_param_groups(self, lr: float, training_params) -> list:
# OPTIMIZE BACKBONE USING CUSTOM LR
backbone_params = {
"named_params": self.backbone.named_parameters(),
"lr": lr * training_params['multiply_backbone_lr'] # You can use any parameter, just make sure to define it when you set up training_params
}
# OPTIMIZE MAIN ARCHITECTURE LAYERS
decoder_named_params = list(self.decoder.named_parameters())
aux_head_named_parameters = list(self.aux_head.named_parameters())
layers_params = {
"named_params": decoder_named_params + aux_head_named_parameters,
"lr": lr
}
param_groups = [backbone_params, layers_params]
return param_groups
def update_param_groups(self, param_groups: list, lr: float, epoch: int, iter: int, training_params, total_batch: int) -> list:
"""
Update the params_groups defined in initialize_param_groups
"""
param_groups[0]["lr"] = lr * training_params['multiply_backbone_lr']
param_groups[1]["lr"] = lr
return param_groups
Note: If working with recipe, don't forget to register your model.
Now you just need to set a value for
multiply_backbone_lr in the training recipe.
yamlmy_training_hyperparams.yaml
...
multiply_backbone_lr: 10 # This is used in our implementation of initialize_param_groups/update_param_groups
optimizer: OptimizerName # Any optimizer as described in the previous sections
optimizer_params: {} # Any parameter for the optimizer you chose
---
Documentation/Source/PhaseCallbacks (documentation/source/PhaseCallbacks.md)
Phase Callbacks
Integrating your own code into an already existing training pipeline can draw much effort on the user's end.
To tackle this challenge, a list of callables triggered at specific points of the training code can
be passed through
training_params.phase_calbacks_list when calling Trainer.train(...).SG's
super_gradients.training.utils.callbacks module implements some common use cases as callbacks: ModelConversionCheckCallback
LRCallbackBase
LinearEpochLRWarmup
LinearBatchLRWarmup
StepLRScheduler
ExponentialLRScheduler
PolyLRScheduler
CosineLRScheduler
FunctionLRScheduler
LRSchedulerCallback
DetectionVisualizationCallback
BinarySegmentationVisualizationCallback
TrainingStageSwitchCallbackBase
YoloXTrainingStageSwitchCallback
For example, the YoloX's COCO detection training recipe uses
YoloXTrainingStageSwitchCallback to turn
off augmentations and incorporate L1 loss starting from epoch 285:super_gradients/recipes/training_hyperparams/coco2017_yolox_train_params.yaml:yamlmax_epochs: 300
...loss: YoloXDetectionLoss
...
phase_callbacks:
- YoloXTrainingStageSwitchCallback:
next_stage_start_epoch: 285
...
Another example is how we use
BinarySegmentationVisualizationCallback to visualize predictions
during training in the Segmentation Transfer Learning Notebook:
How Callbacks work
Callback implements the following methods:text/ Detailed source-code truncated for AI context efficiency. /
The order of the events is as follows:
pythonon_training_start(context) # called once before training starts, good for setting up the warmup LR for epoch in range(epochs):
on_train_loader_start(context)
for batch in train_loader:
on_train_batch_start(context)
on_train_batch_loss_end(context) # called after loss has been computed
on_train_batch_backward_end(context) # called after .backward() was called
on_train_batch_gradient_step_start(context) # called before the optimizer step about to happen (gradient clipping, logging of gradients)
on_train_batch_gradient_step_end(context) # called after gradient step was done, good place to update LR (for step-based schedulers)
on_train_batch_end(context)
on_train_loader_end(context)
on_validation_loader_start(context)
for batch in validation_loader:
on_validation_batch_start(context)
on_validation_batch_end(context)
on_validation_loader_end(context)
on_validation_end_best_epoch(context)
on_test_start(context)
for batch in test_loader:
on_test_batch_start(context)
on_test_batch_end(context)
on_test_end(context)
on_training_end(context) # called once after training ends.
Callbacks are implemented by inheriting this
Callback class, and then by override any of the above-mentioned
method with the wanted behavior. Phase Context
You may have noticed that the
Callback's methods expect a single argument - a PhaseContext instance.PhaseContext includes attributes representing a wide range of training attributes at a given point of the training.text- epoch
- batch_idx
- optimizer
- metrics_dict
- inputs
- preds
- target
- metrics_compute_fn
- loss_avg_meter
- loss_log_items
- criterion
- device
- experiment_name
- ckpt_dir
- net
- lr_warmup_epochs
- sg_logger
- train_loader
- valid_loader
- test_loader
- training_params
- ddp_silent_mode
- checkpoint_params
- architecture
- arch_params
- metric_to_watch
- valid_metrics
- ema_model
- loss_logging_items_names
Each of these attributes is set to
None by default, up until the point it computed or defined in the training pipeline.
- E.g. epoch will be None within on_training_start because, as explained above, this steps happens before the first epoch beginsYou can find which context attribute is set by looking into each method docstring:
pythonclass Callback:
...
def on_training_start(self, context: PhaseContext) -> None:
"""
Called once before start of the first epoch
At this point, the context argument will have the following attributes:
- optimizer
- criterion
- device
- experiment_name
- ckpt_dir
- net
- sg_logger
- train_loader
- valid_loader
- training_params
- checkpoint_params
- arch_params
- metric_to_watch
- valid_metrics The corresponding Phase enum value for this event is Phase.PRE_TRAINING.
:param context:
"""
pass
Build your own Callback
Suppose we would like to implement a simple callback that saves the first batch of images in each epoch for both
training and validation in a new folder called "batch_images" under the local checkpoints directory.
This callback needs to be triggered in 3 places:
1. At the start of training, create a new "batch_images" under the local checkpoints directory.
2. Before passing a train image batch through the network, save it in the new folder.
3. Before passing a validation image batch through the network, save it in the new folder.
Therefore, the callback will override
Callback's on_training_start, on_train_batch_start, and on_validation_batch_start methods:pythonfrom super_gradients.training.utils.callbacks import Callback, PhaseContext
from super_gradients.common.environment.ddp_utils import multi_process_safe
import os
from torchvision.utils import save_image
class SaveFirstBatchCallback(Callback):
def __init__(self):
self.outputs_path = None
self.saved_first_validation_batch = False
@multi_process_safe
def on_training_start(self, context: PhaseContext) -> None:
outputs_path = os.path.join(context.ckpt_dir, "batch_images")
os.makedirs(outputs_path, exist_ok=True)
@multi_process_safe
def on_train_batch_start(self, context: PhaseContext) -> None:
if context.batch_idx == 0:
save_image(context.inputs, os.path.join(self.outputs_path, f"first_train_batch_epoch_{context.epoch}.png"))
@multi_process_safe
def on_validation_batch_start(self, context: PhaseContext) -> None:
if context.batch_idx == 0 and not self.saved_first_validation_batch:
save_image(context.inputs, os.path.join(self.outputs_path, f"first_validation_batch_epoch_{context.epoch}.png"))
self.saved_first_validation_batch = True
IMPORTANTWhen training on multiple nodes (see DDP), the callback will be called at each step once for every
node you are working with. This behaviour may be useful in some specific cases, but in general you will
want to have each method to be triggered only once per step. You can add the decorator
@multi_process_safe to ensure
that only the main node will trigger the callback. In our example, we want to trigger only once per step, so we need to add the
@multi_process_safe decorator.Using Custom Callback within Python Script
The callback can directly be passed through training_params.phase_callbackspythontrainer = Trainer("my_experiment")
train_dataloader = ...
valid_dataloader = ...
model = ...train_params = {
"loss": "CrossEntropyLoss",
"criterion_params": {},
"phase_callbacks": [SaveFirstBatchCallback()],
...
}
trainer.train(training_params=train_params, train_loader=train_dataloader, valid_loader=valid_dataloader)
Using Custom Callback in a Recipe
If you are working with Configuration files, you will be required to do an extra step.
This is similar to using any custom objects in a recipe, and is already defined in the above-mentioned. To summarize, you need to register the new callback by decorating it with the
register_callback decorator,
so that SuperGradients would know how to instantiate it from the .yaml recipe.pythonfrom super_gradients.training.utils.callbacks import Callback, PhaseContext
from super_gradients.common.environment.ddp_utils import multi_process_safe
import os
from torchvision.utils import save_image
from super_gradients.common.registry.registry import register_callback@register_callback()
class SaveFirstBatchCallback(Callback):
def __init__(self):
self.outputs_path = None
self.saved_first_validation_batch = False
@multi_process_safe
def on_training_start(self, context: PhaseContext) -> None:
outputs_path = os.path.join(context.ckpt_dir, "batch_images")
os.makedirs(outputs_path, exist_ok=True)
@multi_process_safe
def on_train_batch_start(self, context: PhaseContext) -> None:
if context.batch_idx == 0:
save_image(context.inputs, os.path.join(self.outputs_path, f"first_train_batch_epoch_{context.epoch}.png"))
@multi_process_safe
def on_validation_batch_start(self, context: PhaseContext) -> None:
if context.batch_idx == 0 and not self.saved_first_validation_batch:
save_image(context.inputs, os.path.join(self.outputs_path, f"first_validation_batch_epoch_{context.epoch}.png"))
self.saved_first_validation_batch = True
Then, in your
my_training_hyperparams.yaml, use SaveFirstBatchCallback in the same way as any other phase callback supported in SG:
yamldefaults:
- default_train_paramsmax_epochs: 250
...
phase_callbacks:
- SaveFirstBatchCallback
Last, make sure to import
SaveFirstBatchCallback in the script you use to launch training from config:
pythonfrom omegaconf import DictConfig
import hydra
import pkg_resources
from my_callbacks import SaveFirstBatchCallback
from super_gradients import Trainer, init_trainer
@hydra.main(config_path=pkg_resources.resource_filename("super_gradients.recipes", ""), version_base="1.2")
def main(cfg: DictConfig) -> None:
Trainer.train_from_config(cfg)
def run():
init_trainer()
main()
if __name__ == "__main__":
run()
This is required, as otherwise
SaveFirstBatchCallback would not be imported at all and therefore SuperGradients
would fail to recognize and instantiate it.---
Documentation/Source/PoseEstimation (documentation/source/PoseEstimation.md)
Pose Estimation
Pose estimation is a computer vision task that involves estimating the position and orientation of objects or people in images or videos.
It typically involves identifying specific keypoints or body parts, such as joints, and determining their relative positions and orientations.
Pose estimation has numerous applications, including robotics, augmented reality, human-computer interaction, and sports analytics.
Top-down and bottom-up are two commonly used approaches in pose estimation. The main difference between top-down and bottom-up pose estimation approaches is the order in which the pose is estimated.
In a top-down approach, an object detection model is used to identify the object of interest, such as a person or a car, and a separate pose estimation model is used to estimate the keypoints of the object.
In contrast, a bottom-up approach first identifies individual body parts or joints and then connects them to form a complete pose.
In summary, top-down approach starts with detecting an object and then estimates its pose, while bottom-up approach first identifies the body parts and then forms a complete pose.
Implemented models
| Model | Model class | Target Generator | Loss Class | Decoding Callback | Visualization Callback |
|------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| DEKR | DEKRPoseEstimationModel | DEKRTargetsGenerator | DEKRLoss | DEKRPoseEstimationDecodeCallback | DEKRVisualizationCallback |
Training
For the sake of being specific in this tutorial, we will consider the training of
DEKR model in further explanations.
The easiest way to start training a pose estimation model is to use a recipe from SuperGradients. Prerequisites
1. You have to install SuperGradients first. Please refer to the Installation section for more details.
2. Prepare the COCO dataset as described in the Computer Vision Datasets Setup under Pose Estimation Datasets
section.
After you met the prerequisites, you can start training the model by running from the root of the repository:
Training from recipe
bashpython -m super_gradients.train_from_recipe --config-name=coco2017_pose_dekr_w32 multi_gpu=Off num_gpus=1
Note, the default configuration for recipe is to use 8 GPUs in DDP mode. This hardware configuration may not be for everyone, so we in the example above we override GPU settings to use single GPU.
It is highly recommended to read through the recipe file to get better understanding of the hyperparameters we use here.
If you're unfamiliar with config files, we recommend you to read the Configuration Files part first.
The start of the config file looks like this:
yamldefaults:
- training_hyperparams: coco2017_dekr_pose_train_params
- dataset_params: coco_pose_estimation_dekr_dataset_params
- arch_params: dekr_w32_arch_params
- checkpoint_params: default_checkpoint_params
- _self_
Here we define the default values for the following parameters:
*
training_hyperparams - These are our training hyperparameters. Things learning rate, optimizer, use of mixed precision, EMA and other training parameters are defined here.
You can refer to the default_train_params.yaml for more details.
In our example we use coco2017_dekr_pose_train_params.yaml that sets
training parameters as in DEKR paper.
* dataset_params - These are the parameters for the training on COCO2017. The dataset configuration sets the dataset transformations (augmentations & preprocessing) and target generator for training the model.
* arch_params - These are the parameters for the model architecture. In our example we use DEKRPoseEstimationModel that is a HRNet-based model with DEKR decoder.
* checkpoint_params - These are the default parameters for resuming of training and using pretrained checkpoints.
You can refer to the default_checkpoint_params.yaml.
Datasets
There are several well-known datasets for pose estimation: COCO, MPII Human Pose, Hands in the Wild, CrowdPose, etc.
SuperGradients provide ready-to-use dataloaders for the COCO dataset COCOKeypointsDataset
and more general
BaseKeypointsDataset implementation that you can subclass from for your specific dataset format.Target generators
The target generators are responsible for generating the target tensors for the model.
Implementation of the target generator is model-specific and usually includes at least a multi-channel heatmap mask per joint.
Each model may require its own target generator implementation that is compatible with model's output.
All target generators should implement
KeypointsTargetsGenerator interface as shown below.
The goal of this class is to transform ground-truth annotations into a format that is suitable for computing a loss and training a model:pysuper_gradients.training.datasets.pose_estimation_datasets.target_generators.KeypointsTargetsGenerator
import abc
import numpy as np
from torch import Tensor
from typing import Union, Tuple, Dict
class KeypointsTargetsGenerator:
@abc.abstractmethod
def __call__(self, image: Tensor, joints: np.ndarray, mask: np.ndarray) -> Union[Tensor, Tuple[Tensor, ...], Dict[str, Tensor]]:
"""
Encode input joints into target tensors
:param image: [C,H,W] Input image tensor
:param joints: [Num Instances, Num Joints, 3] Last channel represents (x, y, visibility)
:param mask: [H,W] Mask representing valid image areas. For instance, in COCO dataset crowd targets
are not used during training and corresponding instances will be zero-masked.
Your implementation may use this mask when generating targets.
:return: Encoded targets
"""
raise NotImplementedError()
SuperGradients provide implementation of DEKRTargetGenerator that is compatible with
DEKR model.If you need to implement your own target generator, please refer to documentation of
KeypointsTargetsGenerator base class. Metrics
A typical metric for pose estimation is the average precision (AP) and average recall (AR).
SuperGradients provide implementation of
PoseEstimationMetrics to compute AP/AR scores.The metric is implemented as a callback that is called after each validation step. Implementation of the metric is made as close as possible to official metric implementation from COCO API.
However, our implementation does NOT include computation of AP/AR scores per area range. It also natively support evaluation in DDP mode.
It is worth noting that usually reported AP/AR scores in papers are obtained using TTA (test-time augmentation) and additional postprocessing on top of the main model.
A horizontal flip is a common TTA technique that is used to increase accuracy of the predictions at the cost of running forward pass twice.
Second common technique is a multi-scale approach when one perform inference additionally on 0.5x and 1.5x input resolution and aggregate predictions.
When training model using SuperGradients, we use neither of these techniques. If you want to measure AP/AR scores using TTA you may want to write your own evaluation loop for that.
In order to use
PoseEstimationMetrics you have to pass a so-called post_prediction_callback to the metric, which is responsible for postprocessing of the model's raw output into final predictions. Postprocessing
Postprocessing refers to a process of transforming the model's raw output into final predictions. Postprocessing is also model-specific and depends on the model's output format.
For
DEKR model, the postprocessing step is implemented in DEKRPoseEstimationDecodeCallback) class.
When instantiating the metric, one has to pass a postprocessing callback as an argument:yamltraining_hyperparams:
valid_metrics_list:
- PoseEstimationMetrics:
num_joints: ${dataset_params.num_joints}
oks_sigmas: ${dataset_params.oks_sigmas}
max_objects_per_image: 20
post_prediction_callback:
_target_: super_gradients.training.utils.pose_estimation.DEKRPoseEstimationDecodeCallback
max_num_people: 20
keypoint_threshold: 0.05
nms_threshold: 0.05
nms_num_threshold: 8
output_stride: 4
apply_sigmoid: False
Visualization
Visualization of the model predictions is a very important part of the training process for pose estimation models.
By visualizing the predicted poses, developers and researchers can identify errors or inaccuracies in the model's output and adjust the model's architecture or training data accordingly.
Overall, visualization is an important tool for improving the accuracy and usability of pose estimation models, both during development and in real-world applications.
SuperGradients provide an implementation of
DEKRVisualizationCallback to visualize predictions for DEKR model.
You can use this callback in your training pipeline to visualize predictions during training. To enable this callback, add the following lines to your training YAML recipe:
yamltraining_hyperparams:
resume: ${resume}
phase_callbacks:
- DEKRVisualizationCallback:
phase:
_target_: super_gradients.training.utils.callbacks.callbacks.Phase
value: TRAIN_BATCH_END
prefix: "train_"
mean: [ 0.485, 0.456, 0.406 ]
std: [ 0.229, 0.224, 0.225 ]
apply_sigmoid: False - DEKRVisualizationCallback:
phase:
_target_: super_gradients.training.utils.callbacks.callbacks.Phase
value: VALIDATION_BATCH_END
prefix: "val_"
mean: [ 0.485, 0.456, 0.406 ]
std: [ 0.229, 0.224, 0.225 ]
apply_sigmoid: False
During training, the callback will generate a visualization of the model predictions and save it to the TensorBoard or Weights & Biases depending on which logger you
are using (Default is Tensorboard). And result will look like this:
On the left side of the image there is input image with ground-truth keypoints overlay and on the right side there are same channel-wise sum of target and predicted heatmaps.
How to connect your own dataset
To add a new dataset to SuperGradients, you need to implement a few things:
- Implement a new dataset class
- Implement a new dataloader factory methods
- Add a configuration file
Let's unwrap each of the steps
Implement a new dataset class
To train an existing architecture on a new dataset one need to implement the dataset class first:
It is generally a good idea to subclass from
BaseKeypointsDataset that gives you a skeleton a dataset class and asks you to implement only a few methods to prepare your data for training.A minimal implementation of a dataset class should look like this:
text/ Detailed source-code truncated for AI context efficiency. /
Implement a new dataloader factory methods
pythonfrom super_gradients.training.dataloaders import get_data_loaderdef my_new_dataset_pose_train(dataset_params: Dict = None, dataloader_params: Dict = None):
return get_data_loader(
config_name="coco_pose_estimation_dataset_params",
dataset_cls=MyNewPoseEstimationDataset,
train=True,
dataset_params=dataset_params,
dataloader_params=dataloader_params,
)
def my_new_dataset_pose_val(dataset_params: Dict = None, dataloader_params: Dict = None):
return get_data_loader(
config_name="coco_pose_estimation_dataset_params",
dataset_cls=MyNewPoseEstimationDataset,
train=False,
dataset_params=dataset_params,
dataloader_params=dataloader_params,
)
Add a configuration file
Create new
my_new_dataset_dataset_params.yaml file under dataset_params folder. For the sake of simplicity, let's assume that we're going to train a DEKR model on human joints (17 keypoints as in COCO).
Then, the full configuration file should look like this:text/ Detailed source-code truncated for AI context efficiency. /
In your training recipe add/change the following lines to:
yamlmy_new_dataset_train_recipe.yaml
defaults:
- training_hyperparams: ...
- dataset_params: my_new_dataset_dataset_params
- arch_params: ...
- checkpoint_params: ...
- _self_
train_dataloader: my_new_dataset_pose_train
val_dataloader: my_new_dataset_pose_val
...
And you should be good to go!
How to add a new model
To implement a new model, you need to add the following parts:
- Model architecture itself
- Target Generator
- Postprocessing Callback
- (Optional) Visualization Callback
A custom target generator class should inherit from
KeypointsTargetsGenerator base class which provides a protocol for generating target tensors for the ground-truth keypoints.
See DEKRTargetsGenerator for more details.A custom postprocessing callback class should have a
forward method which takes raw model predictions and decode them into a final pose predictions.
See DEKRPoseEstimationDecodeCallback for more details.A custom visualization callback class can inherit from
PhaseCallback or Callback base class to generate a visualization of the model predictions.
See DEKRVisualizationCallback for more details.
Rescoring
A rescoring is a third stage of pose estimation (after model forward and nms) aimed to improve the confidence score of the predicted poses.
In a nutshell, rescoring is a multiplication of the final confidence score predicted by the model by a scalar value computed by a rescoring model.
By incorporating the learned prior knowledge about the body structure (in the form for joints linkage information) rescoring model can adjust the final pose confidence
by downweighting the inaccurate of unlikely feasible poses and incresae confidence of poses that are more likely to be correct.
A rescoring model is a simple MLP model that takes the model predictions as tensor of
[B, J, 3] shape as input and outputs a single score for each pose prediction as tensor of [B,1] shape.
Here B represents batch dimension, J number of joints and 3 is the dimension of the joint coordinates (x, y, confidence).To train a rescoring model, you need to have a pretrained pose estimation model first.
SG-TODO: At this point in SG we don't have any pretrained models available. So we should train some models.
Training of rescoring model differs from the regular training in the following ways:
1. Generate the training data.
To train a rescoring model you need to generate the training data first. This assumes that you have a pretrained pose estimation model.
To generate the dataset for rescoding model we run inference on the original dataset (COCO in this example) using our pretrained pose estimation model and save it's predictions to Pickle files.
The rescoring model input are poses
[B,J,3] and the outputs are the rescoring scores [B,1]. The targets are computed object-keypoint similarity (OKs) scores between predicted pose and ground-truth pose.Currently, rescoring is only supported for DEKR architecture.
bashpython -m super_gradients.script.generate_rescoring_training_data --config-name=script_generate_rescoring_data_dekr_coco2017 rescoring_data_dir=OUTPUT_DATA_DIR checkpoint=PATH_TO_TRAINED_MODEL_CHECKPOINT.
2. Train rescoring model.
The training data will be stored in output folder (In the example we use OUTPUT_DATA_DIR placeholder). Once generated you can use this file to train rescoring model:
bashpython -m super_gradients.train_from_recipe --config-name coco2017_pose_dekr_rescoring \
dataset_params.train_dataset_params.pkl_file=OUTPUT_DATA_DIR/rescoring_data_train.pkl \
dataset_params.val_dataset_params.pkl_file=OUTPUT_DATA_DIR/rescoring_data_valid.pkl
This recipe uses custom callback to compute pose estimation metrics on the validation dataset using coordinates of poses from step 1 and confidence values after rescoring.
See integration test case test_dekr_model_with_rescoring for more details and end-to-end usage example.
---
Documentation/Source/PredictionSetup (documentation/source/PredictionSetup.md)
Prediction Set-Up
To make accurate predictions on images, several parameters must be provided:
- Class names: The model predicts class IDs, but to visualize results, the class names from the training dataset are needed.
- Processing parameters: The model requires input data in a specific format.
- Task-specific parameters: For instance, in the case of Detection, this includes IoU and Confidence thresholds.
SuperGradients manages all of these within its model.predict() method, but in certain scenarios, you might need to set these parameters explicitly first.
1. Training your model on a custom dataset
If you trained a model on a dataset that does not inherit from any of the SuperGradients datasets, you will need to set the processing parameters explicitly. To do this, use the model.set_dataset_processing_params() method. Once you've set the parameters, you can run model.predict().2. Using pretrained weights or training on a SuperGradient's dataset
All necessary information is automatically saved during training within the model checkpoint, so you can run model.predict() without calling model.set_dataset_processing_params().For more details about model.predict(), please refer to the related tutorial.
Set-up parameters
Class Names
This is straightforward as it corresponds to the list of classes used during training. For instance, if you're loading the weights of a model fine-tuned on a new dataset, use the classes from that dataset.pythonclass_names = [
"person",
"bicycle",
"car",
"motorcycle",
"airplane",
"bus",
...
]
Ensure that the class order remains the same as during training.Processing
Processing steps are necessary for making predictions.
- Image preprocessing prepares the input data for the model by applying various transformations, such as resizing, normalization, and channel reordering. These transformations ensure the input data is compatible with the model.
- Image postprocessing processes the model's output and converts it into a human-readable and interpretable format. This step may include tasks like converting class probabilities into class labels, applying non-maximum suppression to eliminate duplicate detections, and rescaling results to the original image size.
The super_gradients.training.processing module contains a wide range of Processing transformations responsible for both image preprocessing and postprocessing.
For example, DetectionCenterPadding applies center padding to the image while also handling the reverse transformation to remove padding from the prediction.
Multiple processing transformations can be combined using ComposeProcessing:
pythonfrom super_gradients.training.processing import DetectionCenterPadding, StandardizeImage, NormalizeImage, ImagePermute, ComposeProcessing, DetectionLongestMaxSizeRescaleimage_processor = ComposeProcessing(
[
DetectionLongestMaxSizeRescale(output_shape=(636, 636)),
DetectionCenterPadding(output_shape=(640, 640), pad_value=114),
StandardizeImage(max_value=255.0),
ImagePermute(permutation=(2, 0, 1)),
]
)
Task Specific parameters
#### Detection
Default iou and conf values can be set, which will be used when calling model.predict().
- iou: IoU threshold for the non-maximum suppression (NMS) algorithm. If None, the default value associated with training is used.
- conf: Confidence threshold. Predictions below this threshold are discarded. If None, the default value associated with training is used.
Saving your processing parameters to your model
After defining all parameters, call model.set_dataset_processing_params() and then use model.predict().
pythonfrom super_gradients.common.object_names import Models
from super_gradients.training import modelsmodel = models.get(Models.YOLO_NAS_L, checkpoint_path="/path/to/checkpoint")
model.set_dataset_processing_params(
class_names=class_names,
image_processor=image_processor,
iou=0.35, conf=0.25,
)
IMAGES = [...]
images_predictions = model.predict(IMAGES)
For more information about the model.predict(), please check out the following tutorial.
---
Documentation/Source/Ptq Qat (documentation/source/ptq_qat.md)
Post-Training Quantization (PTQ) and Quantization-Aware Training (QAT)
Content
* Introduction
* Quantization: FP32 vs FP16 vs INT8
* Post-training Quantization
* Quantization-Aware training
* Converting quantized models to ONNX for inference
* Using SuperGradient's Recipes for PTQ/QAT
Introduction
As deep learning models have grown in their complexity and applications,
theyβve also grown large and cumbersome. Large models running on cloud environments have huge compute demand resulting in high cloud cost for developers, posing a major barrier for profitability and scalability. For edge deployments, edge devices are resource-constrained and therefore can not support large and complex models.
Whether the model is deployed on the cloud or at the edge, AI developers are often confronted with the challenge of reducing their model size without compromising model accuracy. Quantization is a common technique used to reduce model size, though it can sometimes result in reduced accuracy.
Quantization aware training is a method that allows practitioners to apply quantization techniques without sacrificing accuracy. It is done in the model training process rather than after the fact. The model size can typically be reduced by two to four times, and sometimes even more.
In this tutorial, weβll compare post-training quantization (PTQ) to quantization-aware training (QAT), and demonstrate how both methods can be easily performed using Deciβs SuperGradients library.
For mode detailed information and theoretical background, refer to this NVIDIA whitepaper and this practical guide from PyTorch.
Note: quantization is currently supported exclusively for GPU and TensorRT environments.
Quantization: FP32 vs FP16 vs INT8
Quantization is a model size reduction technique that converts model weights from high-precision floating-point representation (32-bit float) to low-precision floating-point (FP) representation, such as 16-bit or 8-bit.During quantization, the dynamic range of the original high-precision model has to be compressed into a limited range of the low-precision representation. To achieve this, a calibration process is employed to determine the minimum, maximum, and scale parameters that map the high-precision representation to the low-precision.
The calibration process is performed using a set of representative data samples, known as the calibration dataset, to ensure that the quantization process preserves the model's accuracy as much as possible. The most commonly supported calibration methods are percentile, max, and entropy, which are available in most deep learning frameworks. By using these methods, the quantization process can adapt the parameters based on the specific characteristics of the model and the calibration dataset, resulting in a more accurate quantized model.
Post-training Quantization
Post-training quantization (PTQ) is a quantization method where the quantization process is applied to the trained model after it has completed training. The model's weights and activations are quantized from high precision to low precision, such as from FP32 to INT8. This method is simple and straightforward to implement, but it does not account for the impact of quantization during the training process.
Hybrid quantization
__PREREQUISITE__: You will need pytorch_quantization installed:shellpip install pytorch-quantization --extra-index-url https://pypi.ngc.nvidia.com
With SuperGradients, performing hybrid quantization takes just two lines of code, except of the model definition:pythonimport super_gradients.training.models
from super_gradients.training.utils.quantization.selective_quantization_utils import SelectiveQuantizermodel = super_gradients.training.models.get(model_name="resnet50", pretrained_weights="imagenet")
q_util = SelectiveQuantizer(
default_quant_modules_calibrator_weights="max",
default_quant_modules_calibrator_inputs="histogram",
default_per_channel_quant_weights=True,
default_learn_amax=False,
verbose=True,
)
q_util.quantize_module(model)
Selective quantization
SuperGradients supports selective and partial quantization: skipping modules from quantization, or replacing them with quantization-friendly counterparts. Using the API of SelectiveQuantizer it is straightforward, and it offers great flexibility:
You can skip modules by their names, or by their types:
pythonfrom torch import nnq_util.register_skip_quantization(layer_names={
"layer1",
"layer2.0.conv1",
"conv1"
})
q_util.register_skip_quantization(layer_names={nn.Linear})
You can replace modules with another type, e.g. replace Bottleneck with SuperGradients' QuantBottleneck:
pythonfrom super_gradients.training.models import Bottleneck
from super_gradients.modules.quantization import QuantBottleneckq_util.register_quantization_mapping(layer_names={Bottleneck},
quantized_target_class=QuantBottleneck,
input_quant_descriptor=QuantDescriptor(...),
weights_quant_descriptor=QuantDescriptor(...))
Additionally, if you are designing your own custom block, you can register it, so it will be automatically used for replacement:
pythonfrom super_gradients.training.utils.quantization.selective_quantization_utils import register_quantized_module
from super_gradients.training.utils.quantization.selective_quantization_utils import QuantizedMetadata@register_quantized_module(float_source=MyNonQuantBlock,
action=QuantizedMetadata.ReplacementAction.REPLACE,
input_quant_descriptor=QuantDescriptor(...),
weights_quant_descriptor=QuantDescriptor(...)
)
class MyQuantBlock:
...
#### QuantDescriptor API
QuantDescriptor is a class that is used to configure TensorQuantizer for weights and activations. This class if from pytorch-quantization library and has the following API:
``
Args:
num_bits: An integer. Number of bits of quantization. It is used to calculate scaling factor. Default 8.
name: Seems a nice thing to have
Keyword Arguments:
fake_quant: A boolean. If True, use fake quantization mode. Default True.
axis: None, int or tuple of int. axes which will have its own max for computing scaling factor.
If None (the default), use per tensor scale.
Must be in the range [-rank(input_tensor), rank(input_tensor)).
e.g. For a KCRS weight tensor, quant_axis=(0) will yield per channel scaling.
Default None.
amax: A float or list/ndarray of floats of user specified absolute max range. If supplied,
ignore quant_axis and use this to quantize. If learn_amax is True, will be used to initialize
learnable amax. Default None.
learn_amax: A boolean. If True, learn amax. Default False.
scale_amax: A float. If supplied, multiply amax by scale_amax. Default None. It is useful for some
quick experiment.
calib_method: A string. One of ["max", "histogram"] indicates which calibration to use. Except the simple
max calibration, other methods are all hisogram based. Default "max".
unsigned: A Boolean. If True, use unsigned. Default False.
textUse it to customize your flow. It is recommended to leave default values at least for early experiments.
Quantizing residuals and skip connections
To improve performance of quantized models, quantization of residuals and skip connections is performed. SuperGradients API allows you to do it. In your source code, add one of the following, depending on the type of the skip connection:
python
from super_gradients.modules.skip_connections import (
Residual,
SkipConnection,
CrossModelSkipConnection,
BackboneInternalSkipConnection,
HeadInternalSkipConnection
)
textUse them for all inputs of the sum, mul, div and concat operations. SelectiveQuantizer will take care of them and will replace them with quantized counterparts.
For example, take a simple resnet-like block:
python
from torch import nn
import torch.nn.functional as Fclass ResNetLikeBlock(nn.Module):
def __init__(self, num_channels):
super(ResNetLikeBlock, self).__init__()
self.conv1 = nn.Conv2d(num_channels, num_channels, kernel_size=3, padding=1, bias=False)
self.bn1 = nn.BatchNorm2d(num_channels)
self.conv2 = nn.Conv2d(num_channels, num_channels, kernel_size=3, padding=1, bias=False)
self.bn2 = nn.BatchNorm2d(num_channels)
def forward(self, x):
out = F.relu(self.bn1(self.conv1(x)))
out = F.relu(self.bn2(self.conv2(out)))
out = F.relu(out + x)
return out
textIts quantizeable modification will look like this:
python
from torch import nn
import torch.nn.functional as F
from super_gradients.modules.skip_connections import Residualclass ResNetLikeBlock(nn.Module):
def __init__(self, num_channels):
super(ResNetLikeBlock, self).__init__()
self.conv1 = nn.Conv2d(num_channels, num_channels, kernel_size=3, padding=1, bias=False)
self.bn1 = nn.BatchNorm2d(num_channels)
self.conv2 = nn.Conv2d(num_channels, num_channels, kernel_size=3, padding=1, bias=False)
self.bn2 = nn.BatchNorm2d(num_channels)
self.residual = Residual()
def forward(self, x):
out = F.relu(self.bn1(self.conv1(x)))
out = F.relu(self.bn2(self.conv2(out)))
res = self.residual(x)
out = F.relu(out + res)
return out
textCalibration
And after quantization, performing calibration take another two lines of code:
python
from super_gradients.training.utils.quantization.calibrator import QuantizationCalibrator
model = ... # your quantized model
calib_dataloader = ... # your standard pytorch dataloadercalibrator = QuantizationCalibrator(verbose=True)
calibrator.calibrate_model(
model,
method="percentile",
calib_data_loader=calib_dataloader,
num_calib_batches=16,
percentile=99.99,
)
textYour model is now quantized and calibrated!Refer to
super_gradients/src/super_gradiens/examples/quantization for more source examples that are ready-to-run!
Quantization-Aware training
Quantization-aware training (QAT) is a method that takes into account the impact of quantization during the training process. The model is trained with quantization-aware operations that mimic the quantization process during training. This allows the model to learn how to perform well in the quantized representation, leading to improved accuracy compared to post-training quantization.
With SuperGradients, after you have done PTQ, you can finetune your quantized model with standard training pipeline:
python
from super_gradients import Trainermodel = ... # your quantized and calibrated model
train_dataloader = ... # your standard pytorch dataloader
valid_dataloader = ... # your standard pytorch dataloader
training_hyperparams = ... # refer to training_hyperparams example to fill it
model.train()
trainer = Trainer(experiment_name="my_first_qat_experiment", ckpt_root_dir=...)
res = trainer.train(
model=model,
train_loader=train_dataloader,
valid_loader=valid_dataloader,
training_params=training_hyperparams
)
textAfter that, your model will be finetuned with quantization in mind!Converting quantized models to ONNX for inference
SG is a Production ready library. All the models implemented in SG can be compiled to ONNX, even quantized ones.
If you are using a recipe, neatly quantized ONNX will wait for you in the checkpoints directory.
If you prefer more of a DIY approach, here is the code sample:
python
import torch
from super_gradients.training.utils.quantization.export import export_quantized_module_to_onnxonnx_filename = f"qat_model_1x3x224x224.onnx"
dummy_input = torch.randn([1, 3, 224, 224], device="cpu")
export_quantized_module_to_onnx(
model=quantized_model.cpu(),
onnx_filename=onnx_filename,
input_shape=[1, 3, 224, 224],
input_size=[1, 3, 224, 224],
train=False,
)
textNote that this ONNX uses fake quantization (refer to ONNX QuantizeLinear/DequantizeLinear for more info), while being in FP32 itself. To get a quantized model, you will need an inference framework that will compile ONNX into a runnable engine. Here is an example how to do it with NVIDIA's TensorRT:
shell
trtexec --int8 --fp16 --onnx=qat_model_1x3x224x224.onnx --saveEngine=qat_model_1x3x224x224.pkl
textUsing SuperGradient's Recipes for PTQ/QAT
The SuperGradient library provides a simple and easy-to-use API for both post-training quantization and quantization-aware training. By using the library's recipes, you can quickly and easily quantize models without having to write custom code.
Use
src/super_gradients/examples/qat_from_recipe_example/qat_from_recipe.py to launch your QAT recipes, using train_from_recipe.py will lead you to wrong results!To get a basic understanding of recipes, refer to
configuration_files.md for more details. You can modify an existing recipe to suit PTQ and QAT by adding
quantization_params to it. You can find these default_quantization_params in src/super_gradients/recipes/quantization_params/default_quantization_params.yamlAlso, you can add a sepatare calibration dataloader to your recipe , otherwise, train dataloader without augmenttations will be used for calibration:
yaml
calib_dataloader: imagenet_train # for example
dataset_params:
...
calib_dataloader_params:
...
calib_dataset_params:
...
textInitialization and parameters are identical to training and validation datasets and dataloaders. Refer to configuration_files.md for details.
yaml
ptq_only: False # whether to launch QAT, or leave PTQ only
selective_quantizer_params:
calibrator_w: "max" # calibrator type for weights, acceptable types are ["max", "histogram"]
calibrator_i: "histogram" # calibrator type for inputs acceptable types are ["max", "histogram"]
per_channel: True # per-channel quantization of weights, activations stay per-tensor by default
learn_amax: False # enable learnable amax in all TensorQuantizers using straight-through estimator
skip_modules: # optional list of module names (strings) to skip from quantizationcalib_params:
histogram_calib_method: "percentile" # calibration method for all "histogram" calibrators, acceptable types are ["percentile", "entropy", mse"], "max" calibrators always use "max"
percentile: 99.99 # percentile for all histogram calibrators with method "percentile", other calibrators are not affected
num_calib_batches: # number of batches to use for calibration, if None, 512 / batch_size will be used
verbose: False # if calibrator should be verbose
textAs we have seen earlier, these are the same parameters in the YAML form. If you want to use our rules of thumb to modify your existing training recipe parameters for QAT, you need to use
QATRecipeModificationCallback. To do it, add following config to your recipe:
yaml
pre_launch_callbacks_list:
- QATRecipeModificationCallback:
batch_size_divisor: 2
max_epochs_divisor: 10
lr_decay_factor: 0.01
warmup_epochs_divisor: 10
cosine_final_lr_ratio: 0.01
disable_phase_callbacks: True
disable_augmentations: False
textDefault parameters of this callback are representing the rules of thumb to perform successful QAT from an existing training recipe.---
Documentation/Source/Qat Ptq Yolo Nas (documentation/source/qat_ptq_yolo_nas.md)
PTQ and QAT with YOLO-NAS
<div>
<img src="images/soccer.png" width="750">
</div>In this tutorial, we will guide you step by step on how to prepare our YOLO-NAS for production!
We will leverage YOLO-NAS architecture which includes quantization-friendly blocks, and train a YOLO-NAS model on Roboflow's Soccer Player Detection Dataset
in a way that would maximize our throughput without compromising on the model's accuracy.
The steps will be:
1. Training from scratch on one of the downstream datasets - these will play the role of the user's dataset (i.e., the one in which the model will need to be trained for the user's task)
2. Performing post-training quantization and quantization-aware training
Pre-requisites:
- Training with configuration files
- PTQ and QAT
Note: quantization is currently supported exclusively for GPU and TensorRT environments.
Now, let's get to it.
Step 0: Installations and Dataset Setup
Follow the official instructions to download Roboflow100:
To use this dataset, you must download the "coco" format, NOT the yolov5.
- Your dataset should look like this:
rf100
βββ 4-fold-defect
β ββ train
β β ββ 000000000001.jpg
β β ββ ...
β β ββ _annotations.coco.json
β ββ valid
β β ββ ...
β ββ test
β ββ ...
βββ abdomen-mri
β ββ ...
βββ ...- Install CoCo API: https://github.com/pdollar/coco/tree/master/PythonAPI
textInstall the latest version of SG:
commandline
pip install super-gradients
textInstall torch + PyTorch-quantization (note that later versions should be compatible as well and that you should essentially follow torch installation according to https://pytorch.org/get-started/locally/)
commandline
pip install torch==1.11.0+cu113 torchvision==0.12.0+cu113 torchaudio==0.11.0 --extra-index-url https://download.pytorch.org/whl/cu113 &> /dev/null
pip install pytorch-quantization==2.1.2 --extra-index-url https://pypi.ngc.nvidia.com &> /dev/null
textLaunch Training (non-QA)
Although this might come as a surprise - the name quantization-aware training needs to be more accurate and be performed on a trained checkpoint rather than from scratch.
So in practice, we need to train our model on our dataset fully, then after we perform calibration, we fine-tune our model once again, which will be our final step.
As we discuss in our [Training with configuration files](), we clone the SG repo, then use the repo's configuration files in our training examples.
We will use the
src/super_gradients/recipes/roboflow_yolo_nas_s.yaml`configuration to train the small variant of our DeciModel, DeciModel S.So we navigate to our
`train_from_recipe` script:commandlinecd <YOUR-LOCAL-PATH>/super_gradients/src/super_gradients/examples/train_from_recipe_example
Then to avoid collisions between our cloned and installed SG:
commandlineexport PYTHONPATH=$PYTHONPATH:<YOUR-LOCAL-PATH>/super_gradients/
To launch training on one of the RF100 datasets, we pass it through the dataset_name argument:
text/ Detailed source-code truncated for AI context efficiency. /
And so our best checkpoint resides in <YOUR_CHECKPOINTS_ROOT_DIRECTORY>/yolo_nas_s_soccer_players/ckpt_best.pth reaches 0.967 mAP!
Let's visualize some results:
pythonfrom super_gradients.common.object_names import Models
from super_gradients.training import modelsmodel = models.get(Models.YOLO_NAS_S,
checkpoint_path=<YOUR_CHECKPOINTS_ROOT_DIRECTORY>/yolo_nas_s_soccer_players/ckpt_best.pth>,
num_classes=3)
predictions = model.predict("messi_penalty.mp4")
predictions.show(show_confidence=False)
QAT and PTQ
Now, we will take our checkpoint from our previous section and perform post-training quantization, then quantization-aware training.
To do so, we will need to launch training with our
qat_from_recipe example script, which simplifies taking any existing training recipe and making it a quantization-aware one with the help of some of our recommended practices.
So this time, we navigate to the qat_from_recipe example directory:
commandlinecd <YOUR-LOCAL-PATH>/super_gradients/src/super_gradients/examples/qat_from_recipe_example
Before we launch, let's see how we can easily create a configuration from our
roboflow_yolo_nas_s config to get the most out of QAT and PTQ.
We added a new config that inherits from our previous one, called roboflow_yolo_nas_s_qat.yaml. Let's peek at it:
yamldefaults:
- roboflow_yolo_nas_s
- quantization_params: default_quantization_params
- _self_checkpoint_params:
checkpoint_path: ???
strict_load: no_key_matching
experiment_name: soccer_players_qat_yolo_nas_s
pre_launch_callbacks_list:
- QATRecipeModificationCallback:
batch_size_divisor: 2
max_epochs_divisor: 10
lr_decay_factor: 0.01
warmup_epochs_divisor: 10
cosine_final_lr_ratio: 0.01
disable_phase_callbacks: True
disable_augmentations: False
Let's break it down:
- We inherit from our original non-QA recipe
- We set
quantization_params to the default ones. Reminder - this is where QAT and PTQ hyper-parameters are defined. - We set our checkpoint_params.checkpoint_path to ??? so that passing a checkpoint is required. We will override this value when launching from the command line.
- We add a
QATRecipeModificationCallback to our pre_launch_callbacks_list: This callback accepts the entire cfg: DictConfig and manipulates it right before we start the training. This allows us to adapt any non-QA recipe to a QA one quickly.
Here we will:
- Use half the batch size of the original recipe.
- Use 10 percent of the number of the epochs (and warmup epochs).
- Use 1 percent of the original learning rate.
- Set the final learning rate ratio of the cosine scheduling to 0.01
- Disable augmentations and the phase_callbacks.
Now we can launch PTQ and QAT from the command line:
text/ Detailed source-code truncated for AI context efficiency. /
Observe that for PTQ, our model's mAP decreased from 0.967 to 0.9466. After PTQ, QAT is performed automatically:
text/ Detailed source-code truncated for AI context efficiency. /
We not only observed no decline in the accuracy of our quantized model, but we also gained an improvement of 0.08 mAP!
The QAT model is available in our checkpoints directory, already converted to .onnx format under <YOUR_CHECKPOINTS_ROOT_DIRECTORY>/soccer_players_qat_yolo_nas_s/soccer_players_qat_yolo_nas_s_16x3x640x640_qat.onnx, ready to be converted to converted and deployed to int8 using TRT.
---
Documentation/Source/QuickstartBasicToolkit (documentation/source/QuickstartBasicToolkit.md)
Basic Skills
In this tutorial, we will go over all of the basic functionalities of SuperGradients very briefly.
Go over the following sections to learn how to train, test and predict using SuperGradients.
Check out our extended tutorials on the various features you can find in SuperGradients, and task-specific guides.
<details>
<summary>1. Train a Model</summary>
0. Imports:
pythonfrom super_gradients.common.object_names import Models
from super_gradients.training import Trainer, models
from super_gradients.training.metrics.classification_metrics import Accuracy, Top5
from super_gradients.training.dataloaders.dataloaders import cifar10_train, cifar10_val
from super_gradients.training.utils.distributed_training_utils import setup_device
1. Call
init_trainer() to initialize the super_gradients environment. This should be the first thing to be called by any code running super_gradients:pythoninit_trainer()
2. Call <a href="device.md">setup_device()</a> according to your available hardware and needs. For example, if you want the training to be performed entirely on the CPU:
pythonsetup_device("cpu")
In case multiple GPUs are available, it is also possible to specify the number of GPUs to launch multi-gpu DDP training:
pythonsetup_device(num_gpus=4)
It is also possible to launch the training with whatever available hardware there is (i.e., if there are 4 GPUs available, we will launch a DDP test with four processes) by passing
num_gpus=-1:
pythonsetup_device(num_gpus=-1)
3. Instantiate a Trainer object:
pythontrainer = Trainer(experiment_name="my_cifar_experiment", ckpt_root_dir="/path/to/checkpoints_directory/")
4. <a href="models.md">Instantiate a model</a>:
pythonmodel = models.get(Models.RESNET18, num_classes=10)
5. Define <a href="Metrics.md">metrics</a> and other <a href="https://github.com/Deci-AI/super-gradients/blob/master/src/super_gradients/recipes/training_hyperparams/default_train_params.yaml">training parameters</a>:
pythontraining_params = {
"max_epochs": 20,
"initial_lr": 0.1,
"loss": "CrossEntropyLoss",
"train_metrics_list": [Accuracy(), Top5()],
"valid_metrics_list": [Accuracy(), Top5()],
"metric_to_watch": "Accuracy",
"greater_metric_to_watch_is_better": True,
}
6. Instantiate <a href="https://pytorch.org/tutorials/beginner/basics/data_tutorial.html#preparing-your-data-for-training-with-dataloaders">PyTorch data loaders</a> for training and validation:
pythontrain_loader = cifar10_train()
valid_loader = cifar10_val()
7. Launch training:
pythontrainer.train(model=model, training_params=training_params, train_loader=train_loader, valid_loader=valid_loader)
</details>
<details>
<summary>2. Test a Model</summary>
0. Imports:
pythonfrom super_gradients.common.object_names import Models
from super_gradients.training import Trainer, models
from super_gradients.training.metrics.classification_metrics import Accuracy, Top5
from super_gradients.training.dataloaders.dataloaders import cifar10_val
from super_gradients.training.utils.distributed_training_utils import setup_device
1. Call
init_trainer() to initialize the super_gradients environment. This should be the first thing to be called by any code running super_gradients:pythoninit_trainer()
2. Call <a href="device.md">setup_device()</a> according to your available hardware and needs. For example, if you want the test to be performed entirely on the CPU:
pythonsetup_device("cpu")
In case multiple GPUs are available, it is also possible to specify the number of GPUs to launch a multi-gpu DDP test:
pythonsetup_device(num_gpus=4)
It is also possible to launch the test with whatever available hardware there is (i.e., if there are 4 GPUs available, we will launch a DDP test with four processes) by passing
num_gpus=-1:pythonsetup_device(num_gpus=-1)
3. Instantiate a Trainer object:
pythontrainer = Trainer(experiment_name="test_my_cifar_experiment", ckpt_root_dir="/path/to/checkpoints_directory/")
4. <a href="models.md">Instantiate a model</a> and load weights to it. Learn more about the different options for loading model weights from our <a href="Checkpoints.md">checkpoints tutorial</a>:
pythonmodel = models.get(Models.RESNET18, num_classes=10, checkpoint_path="/path/to/checkpoints_directory/my_cifar_experiment/ckpt_best.pth")
5. Define <a href="Metrics.md">metrics</a> for test:
pythontest_metrics = [Accuracy(), Top5()]
6. Instantiate a <a href="https://pytorch.org/tutorials/beginner/basics/data_tutorial.html#preparing-your-data-for-training-with-dataloaders">PyTorch data loader</a> for testing:
pythontest_data_loader = cifar10_val()
7. Launch test:
pythontest_results = trainer.test(model=model, test_loader=test_data_loader, test_metrics_list=test_metrics)
print(f"Test results: Accuracy: {test_results['Accuracy']}, Top5: {test_results['Top5']}")
</details>
<details>
<summary>3. Use Pre-trained Models</summary>
0. Imports:
pythonfrom super_gradients.common.object_names import Models
from super_gradients.training import models
from super_gradients.training.metrics.classification_metrics import Accuracy, Top5
from super_gradients.training.dataloaders.dataloaders import cifar10_train, cifar10_val
from super_gradients import Trainer, init_trainer
1. Call
init_trainer() to initialize the super_gradients environment. This should be the first thing to be called by any code running super_gradients:pythoninit_trainer()
2. Call <a href="device.md">setup_device()</a> according to your available hardware and needs. For example, if you want the finetuning/test to be performed entirely on the CPU:
pythonsetup_device("cpu")
In case multiple GPUs are available, it is also possible to specify the number of GPUs to launch multi-gpu DDP finetuning/test:
pythonsetup_device(num_gpus=4)
It is also possible to launch the finetuning/test with whatever available hardware there is (i.e., if there are 4 GPUs available, a DDP finetuning/test with four processes will be launched) by passing
num_gpus=-1:pythonsetup_device(num_gpus=-1)
3. Instantiate a pre-trained model from SG's <a href="http://bit.ly/3EGfKD4">model zoo</a>:
pythonmodel = models.get(Models.RESNET18, num_classes=10, pretrained_weights="imagenet")
Or use your local weights to instantiate a pre-trained model:
pythonmodel = models.get(Models.RESNET18, num_classes=10, checkpoint_path="/path/to/imagenet_checkpoint.pth", checkpoint_num_classes=1000)
Finetune or test your pre-trained model as done in the previous sections.
</details>
<details>
<summary>4. Predict</summary>
0. Imports:
pythonfrom PIL import Image
import numpy as np
import requests
from super_gradients.training import models
from super_gradients.common.object_names import Models
import torchvision.transforms as T
import torch
from super_gradients.training.utils.distributed_training_utils import setup_device
1. Call
init_trainer() to initialize the super_gradients environment. This should be the first thing to be called by any code running super_gradients:pythoninit_trainer()
2. Call <a href="device.md">setup_device()</a> according to your available hardware and needs:
pythonsetup_device("cpu")
3. <a href="models.md">Instantiate a model</a>, load weights to it, and put it in
eval mode:pythonLoad the best model that we trained
best_model = models.get(Models.RESNET18, num_classes=10, checkpoint_path="/path/to/checkpoints_directory/my_cifar_experiment/ckpt_best.pth")
best_model.eval()
4. Create input data and preprocess it:
pythonurl = "https://www.aquariumofpacific.org/images/exhibits/Magnificent_Tree_Frog_900.jpg"
image = np.array(Image.open(requests.get(url, stream=True).raw))transforms = T.Compose([
T.ToTensor(),
T.Normalize(mean=(0.4914, 0.4822, 0.4465), std=(0.2023, 0.1994, 0.2010)),
T.Resize((32, 32))
])
input_tensor = transforms(image).unsqueeze(0).to(next(best_model.parameters()).device)
5. Predict and visualize results:
pythonpredictions = best_model(input_tensor)classes = train_dataloader.dataset.classes
plt.xlabel(classes[torch.argmax(predictions)])
plt.imshow(image)
<img src="./images/frog_prediction.png" width="500">
</details>
<details>
<summary>5. Train using SG's Training Recipes</summary>
0. Setup:
- Clone the SG repo:
shellgit clone https://github.com/Deci-AI/super-gradients
- Move to the root of the cloned project (where you find "requirements.txt" and "setup.py") and install super-gradients:
shellpip install -e .
- Append super-gradients to the python path (Replace "YOUR-LOCAL-PATH" with the path to the downloaded repo) to avoid conflicts with any installed version of SG:
shellexport PYTHONPATH=$PYTHONPATH:<YOUR-LOCAL-PATH>/super-gradients/
1. Launch one of SG's <a href="https://github.com/Deci-AI/super-gradients/blob/master/src/super_gradients/recipes">training recipes</a>. For example, Resnet18 on Cifar10:
shellpython -m super_gradients.train_from_recipe --config-name=cifar10_resnet experiment_name=my_resnet18_cifar10_experiment
Learn more in detail on how to launch, customize, and evaluate training recipes from our <a href="configuration_files.md">training with configuration files tutorial</a>.
</details>
---
Documentation/Source/Recipes Custom (documentation/source/Recipes_Custom.md)
Training on Custom Recipes
Prerequisites:
- Introduction to Configuration Files
- Introduction to Training Recipes
- Working with Factories
In this section, we will assume that you want to build you own recipe, and to train a model based on that recipe.
We will cover 2 different approaches in writing your recipe.
1. SuperGradients Format - you stick to the format used in SuperGradients.
2. Custom Format - you organize recipes the way you want.
1. SuperGradient Format
This approach is most appropriate when you want to quickly get started. Since you will be following all SuperGradients convention when building the recipe,
you won't have to worry about working with hydra to instantiate your objects and to launch a training; SuperGradients already provides a script that will do it for you.
How to get started?
1. We recommend that you would go through the pre-defined recipes
and chose the one which seems most similar to your use case. Make sure it covers the same task as you.
2. Copy it to a folder that will be exclusively meant for recipes, inside your project.
3. Override the required parameters to fit your needs. Make sure to keep the same structure. Think about registering custom objects if you need.
4. Copy train_from_recipe script to your project (see below), but think to override
<config-path> with the path to your recipe folder.
pythonThe code below is the same as the
train_from_recipe.py script
See: https://github.com/Deci-AI/super-gradients/blob/master/src/super_gradients/train_from_recipe.py
import hydra
from omegaconf import DictConfig
from super_gradients import Trainer, init_trainer@hydra.main(config_path="<config-path>", version_base="1.2") # TODO: overwrite
<config-path>
def _main(cfg: DictConfig) -> None:
Trainer.train_from_config(cfg)def main() -> None:
init_trainer() #
init_trainer needs to be called before @hydra.main
_main()if __name__ == "__main__":
main()
2. Customizing Recipe Format
With this approach, you will have much more freedom in the way you organize your recipe but this will come at the cost of writing code!
This is mainly recommended for specific use-cases which are not properly covered with the previous approach.
Despite not being required with this approach, we strongly recommend for you to use the same format as in
SuperGradients as it would allow you to build on top of pre-defined recipes.
What are the recipe format constraints here ?
With this approach, you will still need to follow certain conventions
-
training_hyperparams should include the same required fields as with the previous approach. You can find the list here.
- The config passed to dataloaders.get should still be compatible to dataset/dataloader you want to load. Basically, the format constraints that you will face with this approach are the same as these that you would face when working exclusively with python.
How to launch a training ?
Similarly to the previous approach, you will need a script that will launch the training.
The difference being that here you won't be using
Trainer.train_from_config. Instead, you will to isntantiate all the required objects in your script.Here is an example of how such a script could look like:
text/ Detailed source-code truncated for AI context efficiency. /
Tips
Building on top of SuperGradients Recipes
By default, defaults only works with recipes that are defined in the same recipe directory, but this can be extended to other directories.In our case, this comes handy when you want to build on top of recipes that were implemented in SuperGradients.
#### Example
Using
default_train_params defined in super_gradients/recipes/training_hyperparams/default_train_params.yamlyamldefaults:
- training_hyperparams: default_train_params hydra:
searchpath:
- pkg://super_gradients.recipes
... # Continue with your recipe
---
Documentation/Source/Recipes Factories (documentation/source/Recipes_Factories.md)
Working with Factories
Factories in SuperGradients provide a powerful and concise way to instantiate objects in your configuration files.
Prerequisites:
- Training with Configuration Files
- Introduction to Training Recipes
In this tutorial, we'll cover how to use existing factories, register new ones, and briefly explore the implementation details.
Using Existing Factories
If you had a look at the recipes, you may have noticed that many objects are defined directly in the recipes.
In the Supervisely dataset recipe you can see the following
yamltrain_dataset_params:
transforms:
- SegColorJitter:
brightness: 0.1
contrast: 0.1
saturation: 0.1
- SegRandomFlip:
prob: 0.5
- SegRandomRescale:
scales: [0.4, 1.6]
If you load the .yaml recipe as is into a python dictionary, you would get the following
python{
"train_dataset_params": {
"transforms": [
{
"SegColorJitter": {
"brightness": 0.1,
"contrast": 0.1,
"saturation": 0.1
}
},
{
"SegRandomFlip": {
"prob": 0.5
}
},
{
"SegRandomRescale": {
"scales": [0.4, 1.6]
}
}
]
}
}
This configuration alone is not very useful, as we need instances of the classes, not just their configurations.
So we would like to somehow instantiate these classes
SegColorJitter, SegRandomFlip and SegRandomRescale.Factories in SuperGradients come into play here! All these objects were registered beforehand in SuperGradients,
so that when you write these names in the recipe, SuperGradients will detect and instantiate them for you.
Registering a Class
As explained above, only registered objects can be instantiated.
This registration consists of mapping the object name to the corresponding class type.
In the example above, the string
"SegColorJitter" was mapped to the class SegColorJitter, and this is how SuperGradients knows how to convert the string defined in the recipe, into an object.You can register the class using a name different from the actual class name.
However, it's generally recommended to use the same name for consistency and clarity.
Example
pythonfrom super_gradients.common.registry import register_transform@register_transform(name="MyTransformName")
class MyTransform:
def __init__(self, prob: float):
...
In this simple example, we register a new transform.
Note that here we registered (for the sake of the example) the class MyTransform to the name MyTransformName which is different.
We strongly recommend to not do it, and to instead register a class with its own name.Once you registered a class, you can use it in your recipe. Here, we will add this transform to the original recipe
yamltrain_dataset_params:
transforms:
- SegColorJitter:
brightness: 0.1
contrast: 0.1
saturation: 0.1
- SegRandomFlip:
prob: 0.5
- SegRandomRescale:
scales: [0.4, 1.6]
- MyTransformName: # We use the name used to register, which may be different from the name of the class
prob: 0.7
Final Step: Ensure that you import the module containing
MyTransformName into your script.
Doing so will trigger the registration function, allowing SuperGradients to recognize it.Here is an example (adapted from the train_from_recipe script).
pythonfrom .my_module import MyTransform # Importing the module is enough as it will trigger the register_transform functionThe code below is the same as the basic
train_from_recipe.py script
See: https://github.com/Deci-AI/super-gradients/blob/master/src/super_gradients/train_from_recipe.py
from omegaconf import DictConfig
import hydrafrom super_gradients import Trainer, init_trainer
@hydra.main(config_path="recipes", version_base="1.2")
def _main(cfg: DictConfig) -> None:
Trainer.train_from_config(cfg)
def main() -> None:
init_trainer() #
init_trainer needs to be called before @hydra.main
_main()
if __name__ == "__main__":
main()
Under the Hood
Until now, we saw how to use existing Factories, and how to register new ones.
In some cases, you may want to create objects that would benefit from using the factories.
Basic
The basic way to use factories as below.
textfrom super_gradients.common.factories import TransformsFactory
factory = TransformsFactory()
my_transform = factory.get({'MyTransformName': {'prob': 0.7}})
You may recognize that the input passed to factory.get is actually the dictionary that we get after loading the recipe
(See Utilizing Existing Factories)Recommended
Factories become even more powerful when used with the @resolve_param decorator.
This feature allows functions to accept both instantiated objects and their dictionary representations.
It means you can pass either the actual python object or a dictionary that describes it straight from the recipe.pythonclass ImageNetDataset(torch_datasets.ImageFolder):
@resolve_param("transforms", factory=TransformsFactory())
def __init__(self, root: str, transform: Transform):
...
Now,
ImageNetDataset can be passed both an instance of MyTransformpythonmy_transform = MyTransform(prob=0.7)
ImageNetDataset(root=..., transform=my_transform)
And a dictionary representing the same object
pythonmy_transform = {'MyTransformName': {'prob': 0.7}}
ImageNetDataset(root=..., transform=my_transform)
This second way of instantiating the dataset combines perfectly with the concept
.yaml recipes.Difference with
register_transform
- register_transform is responsible to map a string to a class type.
- @resolve_param("transform", factory=TransformsFactory()) is responsible to convert a config into an object, using the mapping created with register_transform. Supported Factory Types
Until here, we focused on a single type of factory, TransformsFactory,
associated with the registration decorator register_transform. SuperGradients supports a wide range of factories, used throughout the training process,
each with its own registering decorator.
SuperGradients offers various types of factories, and each is associated with a specific registration decorator.
` python
from super_gradients.common.factories import (
register_model,
register_kd_model,
register_detection_module,
register_metric,
register_loss,
register_dataloader,
register_callback,
register_transform,
register_dataset,
register_pre_launch_callback,
register_unet_backbone_stage,
register_unet_up_block,
register_target_generator,
register_lr_scheduler,
register_lr_warmup,
register_sg_logger,
register_collate_function,
register_sampler,
register_optimizer,
register_processing,
)
textConclusion
In this tutorial, we have delved into the realm of factories, encompassing:
- Using Existing Factories: How SuperGradients automatically instantiates objects defined in recipes.
- Registering New Classes: The method to map object names to corresponding class types, and how to integrate them in your recipes.
- Under the Hood: Insights into basic and recommended ways to use factories, as well as the variety of supported factory types within SuperGradients.
These insights provide essential understanding and practical techniques to work with factories, a core element in SuperGradients that bridges the gap between configuration and instantiation.
Next Step: Ready to craft your unique recipes? In the next tutorial,
we'll guide you through building your own recipe and training a model based on that recipe.
---
Documentation/Source/Recipes Training (documentation/source/Recipes_Training.md)
Training Recipes
Recipes aim at providing a simple interface to easily reproduce trainings.
Prerequisites
- Introduction to Configuration Files
Training from a Recipe
As explained in our introduction to configuration files, SuperGradients uses the
hydra
library combined with .yaml recipes to allow you to easily customize the parameters.The basic syntax to train a model from a recipe is a follows
bash
python -m super_gradients.train_from_recipe --config-name=<config-name>
textWith <config-name> corresponding to the name of the recipe.You can find all of the pre-defined recipes in super_gradients/recipes.
Recipes usually contain information about their performance, as well as the command to execute them in the header.
Examples
- Training of Resnet18 on Cifar10: super_gradients/recipes/cifar10_resnet.yaml
bash
python -m super_gradients.train_from_recipe --config-name=cifar10_resnet
text- Training of YoloX Small on COCO 2017 (8 GPUs): super_gradients/recipes/coco2017_yolox
bash
python -m super_gradients.train_from_recipe --config-name=coco2017_yolox architecture=yolox_s dataset_params.data_dir=/home/coco2017
textCustomize Training
You may often need to modify certain parameters within a recipe and there are 2 approaches for this:
1. Using hydra overrides.
2. Modifying the recipe.
1. Hydra Overrides
Hydra overrides allow you to change parameters directly from the command line.
This approach is ideal when want to quickly experiment changing a couple of parameters.
Here's the general syntax:
bash
python -m super_gradients.train_from_recipe --config-name=<config-name> param1=<val1> path.to.param2=<val2>
text- Parameters - Listed without the -- prefix.
- Full Path - Use the entire path in the configuration tree, with each level separated by a ..
#### Example
Suppose your recipe looks like this:
yaml
training_hyperparams:
max_epochs: 250
initial_lr: 0.1
...dataset_params:
data_dir: /local/mydataset
...
... # Many other parameters
textChanging Epochs or Learning Rate
bash
python -m super_gradients.train_from_recipe --config-name=<config-name> training_hyperparams.max_epochs=250 training_hyperparams.initial_lr=0.03
textChanging the Dataset Path
bash
python -m super_gradients.train_from_recipe --config-name=<config-name> dataset_params.data_dir=<path-to-dataset>
textNote: Parameter names may differ between recipes, so please check the specific recipe to ensure you're using the correct names.
2. Modifying the Recipe
If you are working on a cloned version of SuperGradients (git clone ...)
then you can directly modify existing recipes. If you installed SuperGradients with pip, then you won't have the ability to modify predefined recipes.
Instead, you should create your own recipe in your project, but you will still have the ability to build it on top of predefined recipes from SuperGradients.
We explain all of this in a following tutorial, but we strongly recommend you to
first finish this tutorial, as it includes information required to fully understand how it works.
Recipe Structure
When browsing the YAML files in the recipes directory, you'll notice that some files contain the key defaults at the beginning of the file.
Here's an example of what this looks like:
yaml
defaults:
- training_hyperparams: cifar10_resnet_train_params
- dataset_params: cifar10_dataset_params
- arch_params: resnet18_cifar_arch_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setuparchitecture: resnet18
train_dataloader: cifar10_train # Optional, see comments below
val_dataloader: cifar10_val # Optional, see comments below
multi_gpu: Off
num_gpus: 1
experiment_suffix: ""
experiment_name: cifar10_${architecture}${experiment_suffix}
textThis is a _minimal_ example of the recipe file that contains all mandatory properties to train a model.Components of a Recipe
We need to introduce some terminology to ensure we stay on the same page throughout the rest of this document.
- Defaults: The
defaults section is critical, and it leverages the OmegaConf syntax. It serves to reference other recipes, allowing you to create modular and reusable configurations.
- Referencing Parameters: This allows you to point to specific parameters in the YAML file according to where they originate. For example, training_hyperparams.initial_lr refers to the initial_lr parameter from the cifar10_resnet_train_params.yaml file.
- Recipe Parameters - _self_: The _self_ keyword has a special role. It permits the current recipe to override the defaults. Its impact depends on its position in the defaults list.A recipe consists of a several sections that are mandatory and required to exist in the recipe file. They are:
-
training_hyperparams - This section contains the hyperparameters related to training regime, such as the learning rate, number of epochs, etc.
- dataset_params - This section contains the parameters related to the dataset and dataloaders for training and validation. Dataset transformations, batch size, etc. are defined here.
The dataset_params section is tightly coupled with the root parameters train_dataloader and val_dataloader.
Please note, that train_dataloader and val_dataloader are optional, not mandatory parameters in a broad sense.
They are used in conjunction to instantiate the dataloaders for training and validation and exists mostly for convenience purposes in SG-provided recipes.
For external datasets we suggest read the Using Custom Datasets section of Datasets documentation page for additional information.
- arch_params - This section contains the parameters related to the model architecture. The arch_params section goes hand-in-hand with the architecture parameter, which is root property of the recipe.
If architecturedefines the specific model architecture, then arch_params defines the parameters for that architecture.
- checkpoint_params - This section contains the parameters related to checkpoints.
It contains settings for loading checkpoint weights for transfer learning, controlling use of pretrained weights and more.
See default_checkpoint_params for an example of what parameters are supported.
- variable_setup: This section is required to enable use of shortcuts for most commonly used overrides which is covered in the next section. Please note it variable_setup must be the last item in the defaults list.Understanding Override Order
π¨ Warning: The order of items in the
defaults section is significant! The overwrite priority follows the list order, meaning that a config defined higher in the list can be overwritten by one defined lower in the list. This is a vital aspect to be aware of when constructing recipes. For a more detailed explanation, please refer to the official documentation.Organizing Your Recipe Folder
Your recipe folder should have a specific structure to match this composition:
ββ cifar10_resnet.yaml
ββ ...
ββtraining_hyperparams
β ββ cifar10_resnet_train_params.yaml
β ββ ...
ββdataset_params
β ββ cifar10_dataset_params.yaml
β ββ ...
ββarch_params
β ββ resnet18_cifar_arch_params.yaml
β ββ ...
ββcheckpoint_params
ββ default_checkpoint_params.yaml
ββ ...
textYou're not restricted to this structure, but following it ensures compatibility with SuperGradients' expectations.Command-Line Override Shortcuts
Although you can override any parameter from the command line, writing the full path of the parameter can be tedious.
For example, to change the learning rate one would have to write
training_hyperparams.initial_lr=0.02.
To change the batch size one would have to write
dataset_params.train_dataloader_params.batch_size=128 dataset_params.val_dataloader_params.batch_size=128.To make it easier, we have defined a few shortcuts for the most common parameters that aims to reduce the amount of typing required:
* Learning rate:
lr=0.02 (same as training_hyperparams.initial_lr=0.02)
* Batch size: bs=128 (same as dataset_params.train_dataloader_params.batch_size=128 dataset_params.val_dataloader_params.batch_size=128)
* Number of train epochs: epochs=100 (same as training_hyperparams.max_epochs=100)
* Number of workers: num_workers=4 (same as dataset_params.train_dataloader_params.num_workers=4 dataset_params.val_dataloader_params.num_workers=4)
* Resume training for a specific experiment: resume=True (same as training_hyperparams.resume=True)
* Enable or disable EMA: ema=true (same as training_hyperparams.ema=true)To use these shortcuts, a
variable_setup section should be a part of hydra defaults in the recipe file.
Please note it variable_setup must be the last item in the defaults list.Conclusion
This tutorial has introduced you to the world of training recipes within SuperGradients. Specifically, you've learned:
- How to Train Models: Utilizing
.yaml recipes to effortlessly train and customize models.
- Ways to Customize Training: Tailoring your training through hydra overrides or direct modifications to the recipes.
- Understanding Recipe Structure: Grasping the organization and conventions that help you align with SuperGradients' expectations.We've laid the groundwork for understanding how recipes enable flexible and reproducible training.
Next Step: In the next tutorial, we'll explore factories in SuperGradients,
revealing how they work with recipes to dynamically instantiate objects. It's a critical step in leveraging the
full power of SuperGradients for your unique needs.
---
Documentation/Source/Segmentation (documentation/source/Segmentation.md)
Image Segmentation
SuperGradients allows users to train models for semantic segmentation tasks.
The library includes pre-trained models, such as the Cityscapes PPLiteSeg model, and provides a simple interface for
loading custom datasets.
Model zoo
SuperGradients includes a variety of pre-trained models for semantic segmentation tasks.
| Model Name | Dataset | IoU | Training Recipe | Resolution |
|----------------|------------|-------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|---------------|
| DDRNet 23 | Cityscapes | 80.26 | cityscapes_ddrnet.yaml | [1024, 2048] |
| DDRNet 23 Slim | Cityscapes | 78.01 | cityscapes_ddrnet.yaml | [1024, 2048] |
| DDRNet 39 | Cityscapes | 81.32 | cityscapes_ddrnet.yaml | [1024, 2048] |
| STDC1 Seg 50 | Cityscapes | 75.11 | cityscapes_stdc_seg50.yaml | [512, 1024] |
| STDC1 Seg 75 | Cityscapes | 76.87 | cityscapes_stdc_seg75.yaml | [768, 1536] |
| STDC2 Seg 50 | Cityscapes | 76.44 | cityscapes_stdc_seg50.yaml | [512, 1024] |
| STDC2 Seg 75 | Cityscapes | 78.93 | cityscapes_stdc_seg75.yaml | [768, 1536] |
| RegSeg 48 | Cityscapes | 78.15 | cityscapes_regseg48.yaml | [1024, 2048] |
| PP-Lite T 50 | Cityscapes | 74.92 | cityscapes_pplite_seg50.yaml | [512, 1024] |
| PP-Lite T 75 | Cityscapes | 77.56 | cityscapes_pplite_seg75.yaml | [512, 1024] |
| PP-Lite B 50 | Cityscapes | 76.48 | cityscapes_pplite_seg50.yaml | [512, 1024] |
| PP-Lite B 75 | Cityscapes | 78.52 | cityscapes_pplite_seg75.yaml | [512, 1024] |
Latency and additional details of these models can be found in the SuperGradients Model Zoo.
Loss functions
SuperGradients provides a variety of loss functions for training semantic segmentation tasks.
All loss functions are implemented in PyTorch and can be found in the
super_gradients.training.losses module.
The following table summarizes the loss functions currently supported by SuperGradients.| Loss function class | Loss name in YAML | Description |
|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------|----------------------------------------------------------------------|
| BCEDiceLoss | bce_dice_loss | Weighted average of BCE and Dice loss |
| CrossEntropyLoss | cross_entropy | Cross entropy loss with label smoothing support |
| DiceLoss | N/A | Dice loss for multiclass segmentation |
| BinaryDiceLoss | N/A | Dice loss for binary segmentation |
| GeneralizedDiceLoss | N/A | Generalized dice loss |
| DiceCEEdgeLoss | dice_ce_edge_loss | Dice loss + Cross entropy loss + Edge loss |
| SegKDLoss | N/A | A loss function for knowledge distillation for semantic segmentation |
Metrics
| Metric Class | Metric name in YAML | Description |
|-------------------------------------------------------------------------------------------------------------------------------------------|---------------------|------------------------------------------------------------------------|
| PixelAccuracy | PixelAccuracy | The ratio of correctly classified pixels to the total number of pixels |
| IoU | IoU | Calculate the Jaccard index for multilabel tasks. |
| Dice | Dice | Calculate the Dice index for multilabel tasks. | |
| Binary IoU | BinaryIOU | Calculate the Jaccard index for binary segmentation task. |
| BinaryDice | BinaryDice | Calculate the Dice index for binary segmentation task. |
See Metrics page for additional details of using metrics in SuperGradients.
Datasets
SuperGradients provides a number of ready to use datasets for semantic segmentation tasks and corresponding data loaders.
| Dataset | Dataset Class | train dataloader | val dataloader |
|------------------|------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|-------------------------------|-----------------------------|
| COCO | CoCoSegmentationDataSet | coco_segmentation_train | coco_segmentation_val |
| Cityscapes | CityscapesDataset | cityscapes_train | cityscapes_val |
| Pascal VOC | PascalVOC2012SegmentationDataSet | pascal_voc_segmentation_train | pascal_voc_segmentation_val |
| Supervisely | SuperviselyPersonDataset | supervisely_persons_train | supervisely_persons_val |
| Mapillary Vistas | MapillaryDataset | mapillary_train | mapillary_val |
In the next section we will demonstrate how to use these datasets and dataloaders to train a segmentation model using SuperGradients.
How to train a segmentation model using Super Gradients
In the tutorial provided, we demonstrate how to fine-tune PPLiteSeg on a subset of the Supervisely dataset.
You can run the following code in our google collab.
Load a dataset
In this example we will work with supervisely-persons. If it's the first time you are using this dataset, or if you want to use another dataset please check out dataset setup instructions
py
from super_gradients.training import dataloadersroot_dir = '/path/to/supervisely_dataset_dir'
train_loader = dataloaders.supervisely_persons_train(dataset_params={"root_dir": root_dir}, dataloader_params={})
valid_loader = dataloaders.supervisely_persons_val(dataset_params={"root_dir": root_dir}, dataloader_params={})
textVisualization
Let's visualize what we've got there.We have images and labels, with the default batch size of 256 for training.
py
from PIL import Image
from torchvision.utils import draw_segmentation_masks
from torchvision.transforms import ToTensor, ToPILImage, Resize
import numpy as np
import torchdef plot_seg_data(img_path: str, target_path: str):
image = (ToTensor()(Image.open(img_path).convert('RGB')) * 255).type(torch.uint8)
target = torch.from_numpy(np.array(Image.open(target_path))).bool()
image = draw_segmentation_masks(image, target, colors="red", alpha=0.4)
image = Resize(size=200)(image)
display(ToPILImage()(image))
for i in range(4, 7):
img_path, target_path = train_loader.dataset.samples_targets_tuples_list[i]
plot_seg_data(img_path, target_path)
textLoad the model from modelzoo
Create a PPLiteSeg nn.Module, with 1 class segmentation head classifier. For simplicity
use_aux_head is set as False
and extra Auxiliary heads aren't used for training.
py
from super_gradients.training import models
from super_gradients.common.object_names import ModelsThe model is a torch.nn.module
model = models.get(
model_name=Models.PP_LITE_T_SEG75, # You can use any model listed in the Models.<Name>
arch_params={"use_aux_heads": False},
num_classes=1, # Change this if you work on another dataset with more classes
pretrained_weights="cityscapes" # Drop this line to train from scratch
)
textNotes- SG includes implementations of
many different architectures.
- Most of these architectures have pretrained checkpoints so feel free to experiment!
- You can use any torch.nn.module model with SuperGradients!
Setup training parameters
The training parameters includes loss, metrics, learning rates and much more. You can check out the default training parameters.
For this task, we will train for 30 epoch, using Binary IoU using the SGD optimizer.
py
from super_gradients.training.metrics.segmentation_metrics import BinaryIOUtrain_params = {
"max_epochs": 30,
"lr_mode": "CosineLRScheduler",
"initial_lr": 0.005,
"lr_warmup_epochs": 5,
"multiply_head_lr": 10,
"optimizer": "SGD",
"loss": "BCEDiceLoss",
"ema": True,
"zero_weight_decay_on_bias_and_bn": True,
"average_best_models": True,
"metric_to_watch": "target_IOU",
"greater_metric_to_watch_is_better": True,
"train_metrics_list": [BinaryIOU()],
"valid_metrics_list": [BinaryIOU()],
"loss_logging_items_names": ["loss"],
}
textLaunch Training
The Trainer in SuperGradient takes care of the entire training and validation process. It serves as a convenient and efficient tool to handle all the details of the training process, allowing you to focus on the development of your model.
py
from super_gradients import Trainertrainer = Trainer(
experiment_name="segmentation_example", # Your experiment checkpoints and logs will be saved in a folder names after the experiment_name.
ckpt_root_dir='/path/to/experiment/folder' # Path to the folder where you want to save all of your experiments.
)
trainer.train(model=model, training_params=training_params, train_loader=train_dataloader, valid_loader=valid_dataloader)
textVisualize the results
py
from torchvision.transforms import Compose, ToTensor, Resize, Normalize, ToPILImagepre_proccess = Compose([
ToTensor(),
Normalize([.485, .456, .406], [.229, .224, .225])
])
demo_img_path = "/home/data/supervisely-persons/images/ache-adult-depression-expression-41253.png"
img = Image.open(demo_img_path)
Resize the image and display
img = Resize(size=(480, 320))(img)
display(img)Run pre-proccess - transforms to tensor and apply normalizations.
img_inp = pre_proccess(img).unsqueeze(0).cuda()Run inference
mask = model(img_inp)Run post-proccess - apply sigmoid to output probabilities, then apply hard
threshold of 0.5 for binary mask prediction.
mask = torch.sigmoid(mask).gt(0.5).squeeze()
mask = ToPILImage()(mask.float())
display(mask)
textGoing further
Troubleshooting
If you encounter any issues, please check out our troubleshooting guide.
How to launch on multiple GPUs (DDP) ?
Please check out our tutorial on how to use multiple GPUs'
How to train models with limited GPU memory?
In case you have a GPU with limited memory, you can use the gradients accumulation technique to "fake" larger batch sizes.
This is not 100% equivalent to training with larger batch sizes, but it is a good approximation.
You can set the desired number of batches to accumulate by changing the
training_hyperparams.batch_accumulate parameter.---
Documentation/Source/SGDocker (documentation/source/SGDocker.md)
SuperGradients Docker Container
Docker is an open-source containerization platform allowing developers to package and distribute applications in a portable and efficient way. Docker is becoming increasingly important in deep learning because it provides an easy and flexible way to manage the complex dependencies and configurations required for deep learning projects. With Docker, deep learning developers can easily package their applications and libraries into container images, which can be distributed and run on any machine with Docker installed. This simplifies the development process and makes it easier to reproduce and share deep learning experiments and results.
Instructions and Recommended Practices
1) Follow the installation steps for the Nvidia Docker.
2) Pull the Docker image with the tag according to the SG version you are working with. For example, super-gradients 3.0.7:
docker pull deciai/super-gradients:3.0.7
textEach SG release will push a new tag to the docker hub.
You can also use the latest tag:
docker pull deciai/super-gradients:latest
textSee the list of available tags here3) Launch the container:
docker run deciai/super-gradients:3.0.7
textRecommendations for training- For the heavier, multi-GPU training, it is best to set the shared memory to at least 64GB by appending
-shm-size=64gb to your run command.
- Add volume mapping for your training data by appending -v /PATH/TO/DATA_DIR/:/PATH/TO/DATA_DIR_INSIDE_THE_CONTAINER/ to your run command. Do the same for your training scripts.
- Make sure all GPUS are accessible by adding --gpus all.
- Run with -it for interactiveness.---
Documentation/Source/Super Gradients.Common (documentation/source/super_gradients.common.rst)
Common package
===============================
.. autosummary::
:toctree: generated
.. automodule:: super_gradients.common
:members:
:undoc-members:
:show-inheritance:
.. automodule:: super_gradients.common.auto_logging
:members:
:undoc-members:
:show-inheritance:
.. automodule:: super_gradients.common.abstraction
:members:
:undoc-members:
:show-inheritance:
.. automodule:: super_gradients.common.data_connection
:members:
:undoc-members:
:show-inheritance:
.. automodule:: super_gradients.common.data_interface
:members:
:undoc-members:
:show-inheritance:
.. automodule:: super_gradients.common.data_types
:members:
:undoc-members:
:show-inheritance:
.. automodule:: super_gradients.common.decorators
:members:
:undoc-members:
:show-inheritance:
.. automodule:: super_gradients.common.environment
:members:
:undoc-members:
:show-inheritance:
.. automodule:: super_gradients.common.factories
:members:
:undoc-members:
:show-inheritance:
.. automodule:: super_gradients.common.plugins
:members:
:undoc-members:
:show-inheritance:
.. automodule:: super_gradients.common.registry
:members:
:undoc-members:
:show-inheritance:
.. automodule:: super_gradients.common.sg_loggers
:members:
:undoc-members:
:show-inheritance:
Module contents
---------------
---
Documentation/Source/Super Gradients (documentation/source/super_gradients.rst)
super\_gradients package
========================
.. toctree::
:maxdepth: 4
super_gradients.common
super_gradients.training
---
Documentation/Source/Super Gradients.Training (documentation/source/super_gradients.training.rst)
Training package
=================================
.. autosummary::
:toctree: generated
.. toctree::
:maxdepth: 4
super_gradients.training
super_gradients.training.dataloaders
super_gradients.training.datasets
super_gradients.training.exceptions
super_gradients.training.kd_trainer
super_gradients.training.legacy
super_gradients.training.losses
super_gradients.training.metrics
super_gradients.training.models
super_gradients.training.sg_trainer
super_gradients.training.training_hyperparams
super_gradients.training.transforms
super_gradients.training.utils
super\_gradients.training module
---------------------------------------
.. automodule:: super_gradients.training
:members:
:undoc-members:
:show-inheritance:
super\_gradients.training.datasets module
---------------------------------------
.. automodule:: super_gradients.training.datasets
:members:
:undoc-members:
:show-inheritance:
super\_gradients.training.dataloaders module
---------------------------------------
.. automodule:: super_gradients.training.dataloaders
:members:
:undoc-members:
:show-inheritance:
super\_gradients.training.exceptions module
---------------------------------------
.. automodule:: super_gradients.training.exceptions
:members:
:undoc-members:
:show-inheritance:
super\_gradients.training.kd_trainer module
---------------------------------------
.. automodule:: super_gradients.training.kd_trainer
:members:
:undoc-members:
:show-inheritance:
super\_gradients.training.legacy module
---------------------------------------
.. automodule:: super_gradients.training.legacy
:members:
:undoc-members:
:show-inheritance:
super\_gradients.training.losses_models module
---------------------------------------------------
.. automodule:: super_gradients.training.losses
:members:
:undoc-members:
:show-inheritance:
super\_gradients.training.metrics module
---------------------------------------------------
.. automodule:: super_gradients.training.metrics
:members:
:undoc-members:
:show-inheritance:
super\_gradients.training.models module
---------------------------------------------------
.. automodule:: super_gradients.training.models
:members:
:undoc-members:
:show-inheritance:
super\_gradients.training.sg\_model module
---------------------------------------------------
.. automodule:: super_gradients.training.sg_trainer
:members:
:undoc-members:
:show-inheritance:
super\_gradients.training.training_hyperparams module
---------------------------------------
.. automodule:: super_gradients.training.training_hyperparams
:members:
:undoc-members:
:show-inheritance:
super\_gradients.training.transforms module
---------------------------------------
.. automodule:: super_gradients.training.transforms
:members:
:undoc-members:
:show-inheritance:
super\_gradients.training.utils module
---------------------------------------------------
.. automodule:: super_gradients.training.utils
:members:
:undoc-members:
:show-inheritance:
Module contents
---------------
---
Documentation/Source/Troubleshooting (documentation/source/troubleshooting.md)
Troubleshooting
This tutorial addresses some of the most frequent concerns we've seen.
If you want more assistance in solving your problem, you may open a new
Issue
in the SuperGradients repository.
CUDA Version error
When using SuperGradients for the first time, you might get this error;
OSError: .../lib/python3.8/site-packages/nvidia/cublas/lib/libcublas.so.11: undefined symbol: cublasLtGetStatusString, version libcublasLt.so.11
textThis may indicate a CUDA conflict between libraries (When Torchvision & Torch are installed for different CUDA versions) or the absence of CUDA support in your Torch version.
To fix this you can- Uninstall both torch and torchvision
pip unistall torch torchvision
- Install the torch version that respects your os & compute platform following the instruction from https://pytorch.org/GPU Memory Overflow
It is pretty common to run out of memory when using GPU. This is shown with following exception:
CUDA out of memory. Tried to allocate 20.00 MiB (GPU 0; 10.76 GiB total capacity; 4.29 GiB already allocated; 10.12 MiB free; 4.46 GiB reserved in total by PyTorch)
textTo reduce memory usage, try the following- Decrease the batch size (
dataset_params.train_dataloader_params.batch_size and dataset_params.val_dataloader_params.batch_size)
- Adjust the number of batch accumulation steps (training_hyperparams.batch_accumulate) and/or number of nodes (if you are using DDP) to keep the effective batch size the same: effective_batch_size = num_gpus batch_size batch_accumulate
CUDA error: device-side assert triggered
You may encounter a generic CUDA error message that lacks information regarding the cause of the error:
RuntimeError: CUDA error: device-side assert triggered
textTo get a better understanding of the root cause of the error, you have the choice between two approaches:1. Run on CPU
When running on CPU you won't have this issue of CUDA hiding the root cause of the error.
2. Set Environment Variable
Some environment variables can be helpful in identifying the root cause:
-
CUDA_LAUNCH_BLOCKING=1 can be used to force synchronous execution of kernel launches, allowing you to pinpoint the exact location of the error in your code.
- CUDA_DEVICE_ASSERT=1 can be used to enable detailed error messages that provide the file name and line number where the assert was triggered.---
Documentation/Source/Welcome (documentation/source/welcome.md)
<div align="center">
<img src="./images/SG - Horizontal Glow 2.png" width="600"/>
<br/><br/>
</div>
SuperGradients
Introduction
Welcome to SuperGradients, a free, open-source training library for PyTorch-based deep learning models.
SuperGradients allows you to train or fine-tune SOTA pre-trained models for all the most commonly applied computer vision tasks with just one training library. We currently support object detection, image classification and semantic segmentation for videos and images.
Why use SuperGradients?
Built-in SOTA Models
Easily load and fine-tune production-ready, pre-trained SOTA models that incorporate best practices and validated hyper-parameters for achieving best-in-class accuracy (Yolox, PP-YoloE, STDC, DDRNet, and PP-LiteSeg).
Easily Reproduce our Results
Why do all the grind work, if we already did it for you? leverage tested and proven recipes & code examples for a wide range of computer vision models generated by our team of deep learning experts. Easily configure your own or use plug & play hyperparameters for training, dataset, and architecture.
Production Readiness and Ease of Integration
All SuperGradients modelsβ are production ready in the sense that they are compatible with deployment tools such as TensorRT (Nvidia) and OpenVINO (Intel) and can be easily taken into production. With a few lines of code you can easily integrate the models into your codebase.
Getting Started
Check out our Quickstart tutorial to get learn the basic of SuperGradients.You can also start from our tutorial on Detection, Segmentation or Pose Estimation.
What's New
__________________________________________________________________________________________________________
Version 3.6.1 (March 6, 2024)
* A dependency from
pycocotools has been removed from SG, we don't rely anymore on this package to parse COCO dataset json.
* A Trainer.ptq and Trainer.qat methods now allow granular control on for the model should be exported (with or without pre-/post-processing).
* A model.predict now has fp16 argument (Default is True) which one can use to disable mixed precision feature (Addressing issues on GTX 16XX series)
* Fixed a bug in missing min-max image normalization in plot() method for detection dataset.
* Removed deci-common from [pro] requirements.
* Updated YoloNAS-Pose fine-tunining for Animals Pose Dataset notebook.
*
__________________________________________________________________________________________________________
Version 3.6.0 (Jan 25, 2024)* Added segmentation samples and support for albumentation transforms for segmentation
* Implemented distance-based detection matching in
DetectionMetrics as an enhancement (by @DimaBir)
* New training hyperparameter - finetune, and multiple LR assignment read about it https://github.com/Deci-AI/super-gradients/blob/master/documentation/source/LRAssignment.md
* Enhanced ImagePermute processing inclusion
* Improved dataset plotting and plot functionality
* A new API for checking model input compatibility
* Extended predict() support for segmentation modelsVersion 3.5.0 (November 23, 2023)
* Support for long videos in
model.predict() (by @hakuryuu96)
* Added support for multiple test loaders in train_from_config
* Added skip_resize to model.predict() to support large images and small objects
Version 3.4.0 (November 6, 2023)* YoloNAS-Pose model released - a new frontier in pose estimation
* Added option to export a recipe to a single YAML file or to a standalone train.py file
* Other bugfixes & minor improvements. Full release notes available here
Citation
If you are using SuperGradients library in your research, please cite SuperGradients deep learning training library.
[//]: # (### BibTeX)
[//]: # ()
[//]: # (
bibtex)[//]: # (@misc{rw2019timm,)
[//]: # ( title = {SuperGradients},)
[//]: # ( year = {2021},)
[//]: # ( publisher = {GitHub},)
[//]: # ( journal = {GitHub repository},)
[//]: # ( doi = {},)
[//]: # ( howpublished = {\url{https://github.com/Deci-AI/super-gradients}})
[//]: # (})
[//]: # (
`)[//]: # ()
[//]: # (### Latest DOI)
[//]: # ()
[//]: # ([](https://zenodo.org/badge/latestdoi/))
Community
If you want to be a part of SuperGradients growing community, hear about all the exciting news and updates, need help, request for advanced features,
or want to file a bug or issue report, we would love to welcome you aboard!
* Slack is the place to be and ask questions about SuperGradients and get support. Click here to join our Slack
* To report a bug, file an issue on GitHub.
* Join the SG Newsletter
for staying up to date with new features and models, important announcements, and upcoming events.
License
This project is released under the Apache 2.0 license.
Citing
BibTeX
bibtex@misc{supergradients,
doi = {10.5281/ZENODO.7789328},
url = {https://zenodo.org/record/7789328},
author = {Aharon, Shay and {Louis-Dupont} and {Ofri Masad} and Yurkova, Kate and {Lotem Fridman} and {Lkdci} and Khvedchenya, Eugene and Rubin, Ran and Bagrov, Natan and Tymchenko, Borys and Keren, Tomer and Zhilko, Alexander and {Eran-Deci}},
title = {Super-Gradients},
publisher = {GitHub},
journal = {GitHub repository},
year = {2021},
}
Latest DOI
[](https://doi.org/10.5281/zenodo.7789328)
---
Documentation/Source/YoloNASPoseQuickstart (documentation/source/YoloNASPoseQuickstart.md)
YOLO-NAS-POSE Quickstart
<div>
<img src="images/yolo_nas_pose_frontier_t4.png" width="750">
</div>Deciβs leveraged its proprietary Neural Architecture Search engine (AutoNAC) to generate YOLO-NAS-POSE - a new object
detection architecture that delivers the worldβs best accuracy-latency performance.
The YOLO-NAS-POSE model incorporates quantization-aware RepVGG blocks to ensure compatibility with post-training
quantization, making it very flexible and usable for different hardware configurations.
In this tutorial, we will go over the basic functionality of the YOLO-NAS-POSE model.
Instantiate a YOLO-NAS-POSE Model
pythonfrom super_gradients.training import models
from super_gradients.common.object_names import Modelsyolo_nas_pose = models.get(Models.YOLO_NAS_POSE_L, pretrained_weights="coco_pose")
Predict
pythonprediction = yolo_nas_pose.predict("https://deci-pretrained-models.s3.amazonaws.com/sample_images/beatles-abbeyroad.jpg")
prediction.show()
<div>
<img src="images/yolo_nas_pose_predict_demo.jpg" width="750">
</div>Export to ONNX & TensorRT
pythonyolo_nas_pose.export("yolo_nas_pose.onnx")
Please follow our Pose Estimation Models Export tutorial for more details.
Evaluation using pycocotools
We provide example notebook to evaluate YOLO-NAS POSE using COCO protocol.
Please check Pose Estimation Models Export tutorial for more details.
---
Documentation/Source/YoloNASQuickstart (documentation/source/YoloNASQuickstart.md)
YOLO-NAS Quickstart
<div>
<img src="images/yolo_nas_frontier.png" width="750">
</div>
Deciβs leveraged its proprietary Neural Architecture Search engine (AutoNAC) to generate YOLO-NAS - a new object detection architecture that delivers the worldβs best accuracy-latency performance. The YOLO-SG model incorporates quantization-aware RepVGG blocks to ensure compatibility with post-training quantization, making it very flexible and usable for different hardware configurations.
In this tutorial, we will go over the basic functionality of the YOLO-NAS model.
Instantiate a YOLO-NAS Model
pythonfrom super_gradients.training import models
from super_gradients.common.object_names import Modelsnet = models.get(Models.YOLO_NAS_S, pretrained_weights="coco")
Predict
pythonprediction = net.predict("https://www.aljazeera.com/wp-content/uploads/2022/12/2022-12-03T205130Z_851430040_UP1EIC31LXSAZ_RTRMADP_3_SOCCER-WORLDCUP-ARG-AUS-REPORT.jpg?w=770&resize=770%2C436&quality=80")
prediction.show()
<div>
<img src="images/yolo_nas_qs_predict.png" width="750">
</div>Export to ONNX
pythonmodels.convert_to_onnx(model=net, input_shape=(3,640,640), out_path="yolo_nas_s.onnx")
Train on RF100
Follow the setup instructions for RF100:
text- Follow the official instructions to download Roboflow100: https://github.com/roboflow/roboflow-100-benchmark?ref=roboflow-blog
//!\\ To use this dataset, you must download the "coco" format, NOT the yolov5. - Your dataset should look like this:
rf100
βββ 4-fold-defect
β ββ train
β β ββ 000000000001.jpg
β β ββ ...
β β ββ _annotations.coco.json
β ββ valid
β β ββ ...
β ββ test
β ββ ...
βββ abdomen-mri
β ββ ...
βββ ...
- Install CoCo API: https://github.com/pdollar/coco/tree/master/PythonAPI
We will use the `roboflow_yolo_nas_s`configuration to train the small variant of our YOLO-NAS, YOLO-NAS-S.To launch training on one of the RF100 datasets, we pass it through the dataset_name argument:
textpython -m super_gradients.train_from_recipe --config-name=roboflow_yolo_nas_s dataset_name=<DATASET_NAME> dataset_params.data_dir=<PATH_TO_RF100_ROOT> ckpt_root_dir=<YOUR_CHECKPOINTS_ROOT_DIRECTORY>
Replace <DATASET_NAME> with any of the RF100 datasets that you wish to train on.
Creating a model for a non-RGB image
You can create a model taking arbitrary number of channels by passing the number of channels to the arch_params argument.
Important thing to keep in mind that in this case you cannot use the available pretrained weights and have to provde
num_classes parameter explicitly.pythonmodel = models.get(Models.YOLO_NAS_S, arch_params=dict(in_channels=2), num_classes=15)
---
Src/Super Gradients/Common/Registry/README (src/super_gradients/common/registry/README.md)
How to use your own objects in SuperGradients recipes ?
1. Introduction
To train a model, it is necessary to configure 4 main components. These components are aggregated into a single "main" recipe .yaml
file that inherits the aforementioned dataset, architecture, training and checkpoint params.
Recipes support out of the box every model, metric or loss that is implemented in SuperGradients, but you can easily extend this to any custom object that you need by "registering it".
Prerequisites
- If you are not familiar with recipes, please check our
- Documentation page on this topic
- Introduction to recipes notebook.*
- All recipes can be found here
2. General flow
In your python script
1. Define your custom object of type:
* metric: torchmetrics.Metric
* model: torch.nn.Module
* loss: torch.nn.modules.loss._Loss
2. Import the associated register decorator:
* metric: from super_gradients.training.utils.registry import register_metric
* model: from super_gradients.training.utils.registry import register_model
* loss: from super_gradients.training.utils.registry import register_loss
* dataloader: from super_gradients.training.utils.registry import register_dataloader
* callback: from super_gradients.training.utils.registry import register_callback
* transform: from super_gradients.training.utils.registry import register_transform
3. Apply it on your object.
* The decorator takes an optional name: str argument. If not specified, the decorated class name will be registered.In your recipe (.yaml)
1. Define your recipe like in any other case (you can find examples here).
2. Modify the recipe by using the registered name (see the following examples).
3. Examples
A. Metric
SuperGradients works with torchmetrics.Metric .
To write your own metric you need to implement update() and compute() methods.In order to work on DDP you also need to define states using add_state().
States are attributes to be reduced, and broadcasted among the different ranks in compute() when training in distributed setting.
An example of state would be the number of correct predictions, which will be summed across the different processes, broadcasted to all of
them before computing the metric value. You can see an example below.
Feel free to check torchmetrics documentation for more information on how to implement your own metric.
main.py
pythonimport omegaconf
import hydraimport torch
import torchmetrics
from super_gradients import Trainer, init_trainer
from super_gradients.common.registry.registry import register_metric
@register_metric() # Will be registered as "CustomTop5"
class CustomTop5(torchmetrics.Metric):
def __init__(self, dist_sync_on_step=False):
super().__init__(dist_sync_on_step=dist_sync_on_step)
self.add_state("correct", default=torch.tensor(0.), dist_reduce_fx="sum")
self.add_state("total", default=torch.tensor(0), dist_reduce_fx="sum")
def update(self, preds: torch.Tensor, target: torch.Tensor):
batch_size = target.size(0)
# Get the top k predictions
_, pred = preds.topk(5, 1, True, True)
pred = pred.t()
# Count the number of correct predictions only for the highest 5
correct = pred.eq(target.view(1, -1).expand_as(pred))
correct5 = correct[:5].reshape(-1).float().sum(0)
self.correct += correct5
self.total += batch_size
def compute(self):
return self.correct.float() / self.total
@hydra.main(config_path="recipes")
def main(cfg: omegaconf.DictConfig) -> None:
Trainer.train_from_config(cfg)
init_trainer()
main()
recipes/training_hyperparams/my_training_hyperparams.yaml
yaml... # Other training hyperparamstrain_metrics_list:
- CustomTop5
valid_metrics_list:
- CustomTop5
Launch the script
bashpython main.py --config-name=my_recipe.yaml
B. Model
pythonimport omegaconf
import hydraimport torch
import torch.nn as nn
import torch.nn.functional as F
from super_gradients import Trainer, init_trainer
from super_gradients.common.registry import register_model
@register_model('my_conv_net') # will be registered as "my_conv_net"
class MyConvNet(nn.Module):
def __init__(self, num_classes: int):
super().__init__()
self.conv1 = nn.Conv2d(3, 6, 5)
self.pool = nn.MaxPool2d(2, 2)
self.conv2 = nn.Conv2d(6, 16, 5)
self.fc1 = nn.Linear(16 5 5, 120)
self.fc2 = nn.Linear(120, 84)
self.fc3 = nn.Linear(84, num_classes)
def forward(self, x):
x = self.pool(F.relu(self.conv1(x)))
x = self.pool(F.relu(self.conv2(x)))
x = torch.flatten(x, 1)
x = F.relu(self.fc1(x))
x = F.relu(self.fc2(x))
x = self.fc3(x)
return x
@hydra.main(config_path="recipes")
def main(cfg: omegaconf.DictConfig) -> None:
Trainer.train_from_config(cfg)
init_trainer()
main()
recipes/my_recipe.yaml
yaml... # Other recipe paramsarchitecture: my_conv_net
Launch the script
bashpython main.py --config-name=my_recipe.yaml
C. Loss
main.py
pythonimport omegaconf
import hydraimport torch
from super_gradients import Trainer, init_trainer
from super_gradients.common.registry.registry import register_loss
@register_loss("custom_rsquared_loss")
class CustomRSquaredLoss(torch.nn.modules.loss._Loss): # The Loss needs to inherit from torch _Loss class.
def forward(self, output, target):
criterion_mse = torch.nn.MSELoss()
return 1 - criterion_mse(output, target).item() / torch.var(target).item()
@hydra.main(config_path="recipes")
def main(cfg: omegaconf.DictConfig) -> None:
Trainer.train_from_config(cfg)
init_trainer()
main()
recipes/training_hyperparams/my_training_hyperparams.yaml
yaml... # Other training hyperparamsloss: custom_rsquared_loss
Launch the script
bashpython main.py --config-name=my_recipe.yaml
---
Src/Super Gradients/Recipes/Cifar10 Resnet.Yaml (src/super_gradients/recipes/cifar10_resnet.yaml)
Cifar10 Classification Training:
Reaches ~94.9 Accuracy after 250 Epochs
Instructions:
0. Make sure that the data is stored in dataset_params.dataset_dir or add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below (feel free to check ReadMe)
1. Move to the project root (where you will find the ReadMe and src folder)
2. Run the command:
python -m super_gradients.train_from_recipe --config-name=cifar10_resnet +experiment_name=cifar10
#
To use equivalent Albumentations transforms pipeline set dataset_params to cifar10_albumentations_dataset_params:
python -m super_gradients.train_from_recipe --config-name=cifar10_resnet dataset_params=cifar10_albumentations_dataset_params
defaults:
- training_hyperparams: cifar10_resnet_train_params
- dataset_params: cifar10_dataset_params
- arch_params: resnet18_cifar_arch_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setuptrain_dataloader: cifar10_train
val_dataloader: cifar10_val
architecture: resnet18_cifar
experiment_name: resnet18_cifar_interpolation_check
multi_gpu: Off
num_gpus: 1
---
Src/Super Gradients/Recipes/Cityscapes Al Ddrnet.Yaml (src/super_gradients/recipes/cityscapes_al_ddrnet.yaml)
Instructions:
0. Make sure that the data is stored in dataset_params.dataset_dir or add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below (feel free to check ReadMe)
1. Prepare the Cityscapes AutoLabelling dataset as described in
src/super_gradients/training/datasets/Dataset_Setup_Instructions.md.
2. Move to the project root (where you will find the ReadMe and src folder)
3. Run the command:
DDRNet39: python -m super_gradients.train_from_recipe --config-name=cityscapes_ddrnet architecture=ddrnet_39
Note: add "checkpoint_params.checkpoint_path=<ddrnet39-backbone-pretrained-path>" to use pretrained backbone
#
Validation mIoU - Cityscapes, training time:
DDRNet39: input-size: [1024, 2048] mIoU: 85.17 4 X RTX A5000, 38 H
#
Pretrained checkpoints:
Backbones- downloaded from the author's official repo.
https://deci-pretrained-models.s3.amazonaws.com/ddrnet/imagenet_pt_backbones/ddrnet39_bb_imagenet.pth
#
Network checkpoints:
DDRNet39: https://sghub.deci.ai/models/ddrnet_39_cityscapes.pth
#
Learning rate and batch size parameters, using 4 RTX A5000 with DDP:
DDRNet39: input-size: [1024, 1024] initial_lr: 0.0075 batch-size: 6 * 4gpus = 24
#
Comments:
* Pretrained backbones were used.
defaults:
- training_hyperparams: cityscapes_default_train_params
- dataset_params: cityscapes_al_dataset_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
architecture: ddrnet_39
dataset_params:
train_dataloader_params:
batch_size: 6
val_dataloader_params:
batch_size: 3
train_dataset_params:
transforms:
- SegColorJitter:
brightness: 0.5
contrast: 0.5
saturation: 0.5
- SegRandomFlip:
prob: 0.5
- SegRandomRescale:
scales: [ 0.5, 2. ]
- SegPadShortToCropSize:
crop_size: [ 1024, 1024 ]
fill_mask: 19
- SegCropImageAndMask:
crop_size: [ 1024, 1024 ]
mode: random
- SegStandardize:
max_value: 255
- SegNormalize:
mean: [0.485, 0.456, 0.406]
std: [0.229, 0.224, 0.225]
- SegConvertToTensor:
mask_output_dtype: long
training_hyperparams:
max_epochs: 200
initial_lr: 0.0075 # batch size 24
loss:
DiceCEEdgeLoss:
num_classes: 19
ignore_index: 19
num_aux_heads: 1
num_detail_heads: 0
weights: [ 1., 0.4 ]
dice_ce_weights: [ 1., 1. ]
ce_edge_weights: [ .5, .5 ]
edge_kernel: 5
sync_bn: True
arch_params:
num_classes: 19
use_aux_heads: True
load_checkpoint: False
checkpoint_params:
load_checkpoint: ${load_checkpoint}
checkpoint_path: ???
load_backbone: True
strict_load: no_key_matching
experiment_name: ${architecture}_cityscapes_al
multi_gpu: DDP
num_gpus: 4
---
Src/Super Gradients/Recipes/Cityscapes Ddrnet.Yaml (src/super_gradients/recipes/cityscapes_ddrnet.yaml)
DDRNet segmentation training example with Cityscapes dataset.
Paper:
"Deep Dual-resolution Networks for Real-time and Accurate Semantic Segmentation of Road Scenes"
https://arxiv.org/abs/2104.13188
#
Instructions:
0. Make sure that the data is stored in dataset_params.dataset_dir or add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below (feel free to check ReadMe)
1. Move to the project root (where you will find the ReadMe and src folder)
2. Run the command:
DDRNet23: python -m super_gradients.train_from_recipe --config-name=cityscapes_ddrnet
DDRNet23-Slim: python -m super_gradients.train_from_recipe --config-name=cityscapes_ddrnet architecture=ddrnet_23_slim
DDRNet39: python -m super_gradients.train_from_recipe --config-name=cityscapes_ddrnet architecture=ddrnet_39
Note: add "checkpoint_params.checkpoint_path=<ddrnet23-backbone-pretrained-path>" to use pretrained backbone
#
Validation mIoU - Cityscapes, training time:
DDRNet23: input-size: [1024, 2048] mIoU: 80.26 4 X RTX A5000, 12 H
DDRNet23-Slim: input-size: [1024, 2048] mIoU: 78.01 4 X RTX A5000, 9 H
DDRNet39: input-size: [1024, 2048] mIoU: 81.32 4 X RTX A5000, 15 H
#
Official git repo:
https://github.com/ydhongHIT/DDRNet
#
Pretrained checkpoints:
Backbones- downloaded from the author's official repo.
https://deci-pretrained-models.s3.amazonaws.com/ddrnet/imagenet_pt_backbones/ddrnet23_bb_imagenet.pth
https://deci-pretrained-models.s3.amazonaws.com/ddrnet/imagenet_pt_backbones/ddrnet23_slim_bb_imagenet.pth
https://deci-pretrained-models.s3.amazonaws.com/ddrnet/imagenet_pt_backbones/ddrnet39_bb_imagenet.pth
#
Logs, tensorboards and network checkpoints:
DDRNet23: https://deci-pretrained-models.s3.amazonaws.com/ddrnet/cityscapes/ddrnet23/
DDRNet23-Slim: https://deci-pretrained-models.s3.amazonaws.com/ddrnet/cityscapes/ddrnet23_slim/
DDRNet39: https://deci-pretrained-models.s3.amazonaws.com/ddrnet/cityscapes/ddrnet39/
#
Learning rate and batch size parameters, using 4 RTX A5000 with DDP:
DDRNet23: input-size: [1024, 1024] initial_lr: 0.0075 batch-size: 6 * 4gpus = 24
DDRNet23-Slim: input-size: [1024, 1024] initial_lr: 0.0075 batch-size: 6 * 4gpus = 24
DDRNet39: input-size: [1024, 1024] initial_lr: 0.0075 batch-size: 6 * 4gpus = 24
#
Comments:
* Pretrained backbones were used.
defaults:
- training_hyperparams: cityscapes_default_train_params
- dataset_params: cityscapes_ddrnet_dataset_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
train_dataloader: cityscapes_train
val_dataloader: cityscapes_val
architecture: ddrnet_23
training_hyperparams:
max_epochs: 500
initial_lr: # batch size 24
default: 0.075
# backbone layers
_backbone: 0.0075
compression3: 0.0075
compression4: 0.0075
down3: 0.0075
down4: 0.0075
layer3_skip: 0.0075
layer4_skip: 0.0075
layer5_skip: 0.0075
loss:
DiceCEEdgeLoss:
num_classes: 19
ignore_index: 19
num_aux_heads: 1
num_detail_heads: 0
weights: [ 1., 0.4 ]
dice_ce_weights: [ 1., 1. ]
ce_edge_weights: [ .5, .5 ]
edge_kernel: 5
sync_bn: True
arch_params:
num_classes: 19
use_aux_heads: True
load_checkpoint: False
checkpoint_params:
load_checkpoint: ${load_checkpoint}
checkpoint_path:
load_backbone: True
strict_load: no_key_matching
experiment_name: ${architecture}_cityscapes
multi_gpu: DDP
num_gpus: 4
---
Src/Super Gradients/Recipes/Cityscapes Kd Base.Yaml (src/super_gradients/recipes/cityscapes_kd_base.yaml)
Distillation for semantic segmentation on Cityscapes dataset.
#
Instructions:
0. Make sure that the data is stored in dataset_params.[train/val]_dataset_params.root_dir or
add "dataset_params.[train/val]_dataset_params.root_dir=<PATH-TO-DATASET>" at the end of the
command below (feel free to check ReadMe)
1. Move to the project root (where you will find the ReadMe and src folder)
2. Run the command:
DDRNet23: python -m super_gradients.train_from_kd_recipe --config-name=cityscapes_kd_base student_architecture=ddrnet_23
DDRNet23-Slim: python -m super_gradients.train_from_kd_recipe --config-name=cityscapes_kd_base student_architecture=ddrnet_23_slim
Note: add "student_checkpoint_params.checkpoint_path=<ddrnet23-backbone-pretrained-path>" to use pretrained backbone
#
Teachers specifications:
DDRNet39-AL: mIoU: 85.17 notes: trained with Cityscapes coarse data.
#
Validation mIoU results - Cityscapes, training time:
DDRNet23: teacher: DDRNet39-AL input-size: [1024, 2048] mIoU: 81.48 4 X RTX A5000, 13 H
DDRNet23-Slim: teacher: DDRNet39-AL input-size: [1024, 2048] mIoU: 79.41 4 X RTX A5000, 11 H
#
Pretrained backbones checkpoints:
https://deci-pretrained-models.s3.amazonaws.com/ddrnet/imagenet_pt_backbones/ddrnet23_bb_imagenet.pth
https://deci-pretrained-models.s3.amazonaws.com/ddrnet/imagenet_pt_backbones/ddrnet23_slim_bb_imagenet.pth
#
Logs, tensorboards and network checkpoints:
DDRNet23: https://deci-pretrained-models.s3.amazonaws.com/ddrnet/cityscapes/ddrnet23_cwd/
DDRNet23-Slim: https://deci-pretrained-models.s3.amazonaws.com/ddrnet/cityscapes/ddrnet23_slim_cwd/
#
Learning rate and batch size parameters, using 4 RTX A5000 with DDP:
DDRNet23: input-size: [1024, 1024] initial_lr: 0.0075 batch-size: 6 * 4gpus = 24
DDRNet23-Slim: input-size: [1024, 1024] initial_lr: 0.0075 batch-size: 6 * 4gpus = 24
#
Teachers checkpoints:
DDRNet39-AL: https://deci-pretrained-models.s3.amazonaws.com/ddrnet/cityscapes/ddrnet39_al/average_model_2023_02_20.pth
#
Comments:
* Pretrained backbones were used for the student models.
* Default hyper-parameters are based on DDRNet model train recipes, for full resolution training [1024 x 2048]
defaults:
- training_hyperparams: cityscapes_default_train_params
- dataset_params: cityscapes_ddrnet_dataset_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
train_dataloader: cityscapes_train
val_dataloader: cityscapes_val
resume: False
training_hyperparams:
sync_bn: True
max_epochs: 500
initial_lr: # batch size 24
default: 0.075
# backbone layers
_backbone: 0.0075
compression3: 0.0075
compression4: 0.0075
down3: 0.0075
down4: 0.0075
layer3_skip: 0.0075
layer4_skip: 0.0075
layer5_skip: 0.0075
resume: ${resume}
loss:
_target_: super_gradients.training.losses.seg_kd_loss.SegKDLoss
weights: [ 1. ]
kd_loss_weights: [1., 6.]
kd_loss:
_target_: super_gradients.training.losses.cwd_loss.ChannelWiseKnowledgeDistillationLoss
temperature: 3.
normalization_mode: channel_wise
ce_loss:
_target_: torch.nn.CrossEntropyLoss
ignore_index: 19
student_arch_params:
num_classes: 19
use_aux_heads: False
teacher_arch_params:
num_classes: 19
use_aux_heads: False
KD module arch params
arch_params:teacher_checkpoint_params:
load_backbone:
checkpoint_path:
strict_load: no_key_matching
pretrained_weights: cityscapes
student_checkpoint_params:
load_backbone: True
checkpoint_path: ??? # ImageNet pretrained checkpoints
strict_load: no_key_matching
pretrained_weights:
run_teacher_on_eval: True
multi_gpu: DDP
num_gpus: 4
architecture: kd_module
student_architecture: ???
teacher_architecture: ddrnet_39
experiment_name: ${student_architecture}_teacher-${teacher_architecture}
---
Src/Super Gradients/Recipes/Cityscapes Pplite Seg50.Yaml (src/super_gradients/recipes/cityscapes_pplite_seg50.yaml)
PPLiteSeg segmentation training example with Cityscapes dataset.
Torch implementation of the paper:
Juncai Peng, Yi Liu, Shiyu Tang, Yuying Hao, Lutao Chu, Guowei Chen, Zewu Wu, Zeyu Chen, Zhiliang Yu, Yuning Du,
Qingqing Dang,Baohua Lai, Qiwen Liu, Xiaoguang Hu, Dianhai Yu, Yanjun Ma.
PP-LiteSeg: A Superior Real-Time Semantic Segmentation Model.
Instructions:
0. Make sure that the data is stored in dataset_params.dataset_dir or add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below (feel free to check ReadMe)
1. Move to the project root (where you will find the ReadMe and src folder)
2. Run the command:
PPLite-T-Seg50: python -m super_gradients.train_from_recipe --config-name=cityscapes_pplite_seg50 checkpoint_params.checkpoint_path=<stdc1-backbone-pretrained-path> architecture=pp_lite_t_seg
PPLite-B-Seg50: python -m super_gradients.train_from_recipe --config-name=cityscapes_pplite_seg50 checkpoint_params.checkpoint_path=<stdc2-backbone-pretrained-path> architecture=pp_lite_b_seg
#
#
Validation mIoU - Cityscapes, training time:
PPLite-T-Seg50: input-size: [512, 1024] mIoU: 74.92 4 X RTX A5000, 13 H
PPLite-B-Seg50: input-size: [512, 1024] mIoU: 76.48 4 X RTX A5000, 14 H
#
Official git repo:
https://github.com/PaddlePaddle/PaddleSeg/
Paper:
https://arxiv.org/abs/2204.02681
#
Pretrained checkpoints:
Backbones- downloaded from the STDC author's official repo.
PPLite-T-Seg50, (STDC1-backbone): https://deci-pretrained-models.s3.amazonaws.com/stdc_backbones/stdc1_imagenet_pretrained.pth
PPLite-B-Seg50, (STDC2-backbone): https://deci-pretrained-models.s3.amazonaws.com/stdc_backbones/stdc2_imagenet_pretrained.pth
#
Logs, tensorboards and network checkpoints:
PPLite-T-Seg50: https://deci-pretrained-models.s3.amazonaws.com/ppliteseg/cityscapes/pplite_t_seg50/
PPLite-B-Seg50: https://deci-pretrained-models.s3.amazonaws.com/ppliteseg/cityscapes/pplite_b_seg50/
#
Learning rate and batch size parameters, using 2 RTX A5000 with DDP:
PPLite-T-Seg50: input-size: [512, 1024] initial_lr: 0.01 batch-size: 8 * 4gpus = 32
PPLite-B-Seg50: input-size: [512, 1024] initial_lr: 0.01 batch-size: 8 * 4gpus = 32
#
Comments:
* ImageNet Pretrained backbones were used.
defaults:
- training_hyperparams: cityscapes_default_train_params
- dataset_params: cityscapes_stdc_seg50_dataset_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
train_dataloader: cityscapes_train
val_dataloader: cityscapes_val
architecture: pp_lite_t_seg
dataset_params:
train_dataloader_params:
batch_size: 8
val_dataloader_params:
batch_size: 8
arch_params:
num_classes: 19
use_aux_heads: True
checkpoint_params:
checkpoint_path:
load_backbone: True
load_weights_only: True
strict_load: no_key_matching
training_hyperparams:
sync_bn: True
initial_lr:
"encoder.backbone": 0.01
default: 0.1
loss:
DiceCEEdgeLoss:
num_classes: 19
ignore_index: 19
num_aux_heads: 3
num_detail_heads: 0
weights: [ 1., 1., 1., 1. ]
dice_ce_weights: [ 1., 1. ]
ce_edge_weights: [ .5, .5 ]
edge_kernel: 5
multi_gpu: DDP
num_gpus: 4
experiment_name: ${architecture}50_cityscapes
---
Src/Super Gradients/Recipes/Cityscapes Pplite Seg75.Yaml (src/super_gradients/recipes/cityscapes_pplite_seg75.yaml)
PPLiteSeg segmentation training example with Cityscapes dataset.
Torch implementation of the paper:
Juncai Peng, Yi Liu, Shiyu Tang, Yuying Hao, Lutao Chu, Guowei Chen, Zewu Wu, Zeyu Chen, Zhiliang Yu, Yuning Du,
Qingqing Dang,Baohua Lai, Qiwen Liu, Xiaoguang Hu, Dianhai Yu, Yanjun Ma.
PP-LiteSeg: A Superior Real-Time Semantic Segmentation Model.
Instructions:
0. Make sure that the data is stored in dataset_params.dataset_dir or add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below (feel free to check ReadMe)
1. Move to the project root (where you will find the ReadMe and src folder)
2. Run the command:
PPLite-T-Seg75: python -m super_gradients.train_from_recipe --config-name=cityscapes_pplite_seg75 checkpoint_params.checkpoint_path=<stdc1-backbone-pretrained-path> architecture=pp_lite_t_seg
PPLite-B-Seg75: python -m super_gradients.train_from_recipe --config-name=cityscapes_pplite_seg75 checkpoint_params.checkpoint_path=<stdc2-backbone-pretrained-path> architecture=pp_lite_b_seg
#
#
Validation mIoU - Cityscapes, training time:
PPLite-T-Seg75: input-size: [768, 1536] mIoU: 77.56 4 X RTX A5000, 13 H
PPLite-B-Seg75: input-size: [768, 1536] mIoU: 78.52 4 X RTX A5000, 14 H
#
Official git repo:
https://github.com/PaddlePaddle/PaddleSeg/
Paper:
https://arxiv.org/abs/2204.02681
#
Pretrained checkpoints:
Backbones- downloaded from the STDC author's official repo.
PPLite-T-Seg75, (STDC1-backbone): https://deci-pretrained-models.s3.amazonaws.com/stdc_backbones/stdc1_imagenet_pretrained.pth
PPLite-B-Seg75, (STDC2-backbone): https://deci-pretrained-models.s3.amazonaws.com/stdc_backbones/stdc2_imagenet_pretrained.pth
#
Logs, tensorboards and network checkpoints:
PPLite-T-Seg75: https://deci-pretrained-models.s3.amazonaws.com/ppliteseg/cityscapes/pplite_t_seg75/
PPLite-B-Seg75: https://deci-pretrained-models.s3.amazonaws.com/ppliteseg/cityscapes/pplite_b_seg75/
#
Learning rate and batch size parameters, using 2 RTX A5000 with DDP:
PPLite-T-Seg75: input-size: [768, 768] initial_lr: 0.01 batch-size: 8 * 4gpus = 32
PPLite-B-Seg75: input-size: [768, 768] initial_lr: 0.01 batch-size: 8 * 4gpus = 32
#
Comments:
* ImageNet Pretrained backbones were used.
defaults:
- training_hyperparams: cityscapes_default_train_params
- dataset_params: cityscapes_ppliteseg_seg75_dataset_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
train_dataloader: cityscapes_train
val_dataloader: cityscapes_val
architecture: pp_lite_t_seg
arch_params:
num_classes: 19
use_aux_heads: True
checkpoint_params:
checkpoint_path:
load_backbone: True
load_weights_only: True
strict_load: no_key_matching
training_hyperparams:
sync_bn: True
initial_lr:
"encoder.backbone": 0.01
default: 0.1
loss:
DiceCEEdgeLoss:
num_classes: 19
ignore_index: 19
num_aux_heads: 3
num_detail_heads: 0
weights: [ 1., 1., 1., 1. ]
dice_ce_weights: [ 1., 1. ]
ce_edge_weights: [ .5, .5 ]
edge_kernel: 5
multi_gpu: DDP
num_gpus: 4
experiment_name: ${architecture}75_cityscapes
---
Src/Super Gradients/Recipes/Cityscapes Regseg48.Yaml (src/super_gradients/recipes/cityscapes_regseg48.yaml)
RegSeg segmentation training example with Cityscapes dataset.
Reproduction of paper: Rethink Dilated Convolution for Real-time Semantic Segmentation.
#Instructions:
0. Make sure that the data is stored in dataset_params.dataset_dir or add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below (feel free to check ReadMe)
1. Move to the project root (where you will find the ReadMe and src folder)
2. Run the command:
python -m super_gradients.train_from_recipe --config-name=cityscapes_regseg48
#
#
Validation mIoU - Cityscapes, training time:
RegSeg48: input-size: [1024, 2048] mIoU: 78.15 using 4 GeForce RTX 2080 Ti with DDP, ~2 minutes / epoch
#
Official git repo:
https://github.com/RolandGao/RegSeg
Paper:
https://arxiv.org/pdf/2111.09957.pdf
#
#
Logs, tensorboards and network checkpoints:
s3://deci-pretrained-models/regseg48_cityscapes/
#
#
Learning rate and batch size parameters, using 4 GeForce RTX 2080 Ti with DDP:
RegSeg48: input-size: [1024, 2048] initial_lr: 0.02 batch-size: 4 * 4gpus = 16
defaults:
- training_hyperparams: default_train_params
- dataset_params: cityscapes_regseg48_dataset_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
train_dataloader: cityscapes_train
val_dataloader: cityscapes_val
cityscapes_ignored_label: 19 # convenience parameter since it is used in many places in the YAML
architecture: regseg48
arch_params:
num_classes: 19
strict_load: no_key_matching
load_checkpoint: False
resume: False
training_hyperparams:
sync_bn: True
resume: ${resume}
max_epochs: 800
lr_mode: PolyLRScheduler
initial_lr: 0.02 # for effective batch_size=16
lr_warmup_epochs: 0
optimizer: SGD
optimizer_params:
momentum: 0.9
weight_decay: 5e-4
ema: True
loss: CrossEntropyLoss
criterion_params:
ignore_index: ${cityscapes_ignored_label}
train_metrics_list:
- PixelAccuracy:
ignore_label: ${cityscapes_ignored_label}
- IoU:
num_classes: 20
ignore_index: ${cityscapes_ignored_label}
valid_metrics_list:
- PixelAccuracy:
ignore_label: ${cityscapes_ignored_label}
- IoU:
num_classes: 20
ignore_index: ${cityscapes_ignored_label}
metric_to_watch: IoU
greater_metric_to_watch_is_better: True
_convert_: all
project_name: RegSeg
experiment_name: ${architecture}_cityscapes
multi_gpu: AUTO
num_gpus: 4
---
Src/Super Gradients/Recipes/Cityscapes Segformer.Yaml (src/super_gradients/recipes/cityscapes_segformer.yaml)
SegFormer segmentation training example with Cityscapes dataset.
Reproduction of paper:
Enze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar, Jose M. Alvarez, Ping Luo
"SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers"
( https://arxiv.org/pdf/2105.15203.pdf )
#
Official git repo:
https://github.com/NVlabs/SegFormer
#
Code and Imagenet-1k pre-trained backbone weights taken and adapted from:
https://github.com/sithu31296/semantic-segmentation
#
Instructions:
In the recipe of the specif variant you would like to train:
1. Choose SegFormer architecture (b0 - b5) by changing the value of the "architecture".
2. We recommend preparing the data according to SG's CityScapes readme file:
https://github.com/Deci-AI/super-gradients/blob/master/src/super_gradients/training/datasets/Dataset_Setup_Instructions.md
3. Note: if you change the dataset's internal directory structure, make changes to the fields "list_file" and
"labels_csv_path" of both "train_dataset_params" and "val_dataset_params" accordingly
4. Edit the "data_root_dir" field to point to the absolute path of the data root directory
5. Edit the "ckpt_root_dir" field to the path where you want to save checkpoints and logs
6. Move to the project root (where you will find the ReadMe and src folder)
7. Run the command (change the config_name according to the variant):
python -m super_gradients.train_from_recipe --config-name=cityscapes_segformer_b0
#
#
Imagenet-1K pre-trained backbone:
MiT (Mix Transformer) B0: https://deci-pretrained-models.s3.amazonaws.com/mit_backbones/mit_b0.pth
B1: https://deci-pretrained-models.s3.amazonaws.com/mit_backbones/mit_b1.pth
B2: https://deci-pretrained-models.s3.amazonaws.com/mit_backbones/mit_b2.pth
B3: https://deci-pretrained-models.s3.amazonaws.com/mit_backbones/mit_b3.pth
B4: https://deci-pretrained-models.s3.amazonaws.com/mit_backbones/mit_b4.pth
B5: https://deci-pretrained-models.s3.amazonaws.com/mit_backbones/mit_b5.pth
#
8. Download the weights from the above link and put them in a directory of your choice
9. Insert the weights file's full path to checkpoint_params.checkpoint_path:
10. Ensure checkpoint_params.load_backbone: True
#
#
Performance and training details:
SegFormer-B0: mIoU (sliding-window inference) on validation set: 76.14
training time: 6 hours with 8 NVIDIA RTX A5000 GPUs with DDP, ~1 minutes / epoch
SegFormer-B1: mIoU (sliding-window inference) on validation set: 77.80
training time: 8 hours with 8 NVIDIA RTX A5000 GPUs with DDP, ~1 minutes / epoch
SegFormer-B2: mIoU (sliding-window inference) on validation set: 81.43
training time: 26 hours with 4 NVIDIA RTX A5000 GPUs with DDP, ~4 minutes / epoch
SegFormer-B3: mIoU (sliding-window inference) on validation set: 82.24
training time: 36 hours with 4 NVIDIA RTX A5000 GPUs with DDP, ~3 minutes / epoch
SegFormer-B4: mIoU (sliding-window inference) on validation set: 82.39
training time: 25 hours with 4 NVIDIA RTX A5000 GPUs with DDP, ~3 minutes / epoch
SegFormer-B5: mIoU (sliding-window inference) on validation set: 82.33
training time: 41 hours with 4 NVIDIA RTX A5000 GPUs with DDP, ~3 minutes / epoch
defaults:
- training_hyperparams: default_train_params
- dataset_params: cityscapes_segformer_dataset_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
data_root_dir: /data/cityscapes
dataset_params:
train_dataset_params:
root_dir: ${data_root_dir}
val_dataset_params:
root_dir: ${data_root_dir}
experiment_name: ${architecture}_cityscapes
train_dataloader: cityscapes_train
val_dataloader: cityscapes_val
cityscapes_ignored_label: 19 # convenience parameter since it is used in many places in the YAML
arch_params:
num_classes: 19
checkpoint_params:
checkpoint_path:
load_backbone: True
load_weights_only: True
strict_load: no_key_matching
load_checkpoint: False
resume: False
training_hyperparams:
resume: ${resume}
optimizer: AdamW
zero_weight_decay_on_bias_and_bn: True
sync_bn: True
loss: CrossEntropyLoss
criterion_params:
ignore_index: ${cityscapes_ignored_label}
phase_callbacks:
- SlidingWindowValidationCallback:
transforms_for_sliding_window: []
train_metrics_list:
- IoU:
num_classes: 20
ignore_index: ${cityscapes_ignored_label}
valid_metrics_list:
- IoU:
num_classes: 20
ignore_index: ${cityscapes_ignored_label}
metric_to_watch: IoU
greater_metric_to_watch_is_better: True
---
Src/Super Gradients/Recipes/Cityscapes Segformer B0.Yaml (src/super_gradients/recipes/cityscapes_segformer_b0.yaml)
defaults:
- cityscapes_segformer
- _self_
- variable_setup
architecture: segformer_b0
checkpoint_params:
checkpoint_path:
load_backbone: True
load_weights_only: True
strict_load: no_key_matching
load_checkpoint: False
experiment_name: ${architecture}_cityscapes
training_hyperparams:
max_epochs: 2
lr_mode: PolyLRScheduler
initial_lr: 0.00006 # for effective batch_size=8
multi_gpu: DDP
num_gpus: 4
---
Src/Super Gradients/Recipes/Cityscapes Segformer B1.Yaml (src/super_gradients/recipes/cityscapes_segformer_b1.yaml)
defaults:
- cityscapes_segformer
- _self_
- variable_setup
architecture: segformer_b1
checkpoint_params:
checkpoint_path:
load_backbone: True
load_weights_only: True
strict_load: no_key_matching
load_checkpoint: False
experiment_name: ${architecture}_cityscapes
training_hyperparams:
max_epochs: 2
lr_mode: PolyLRScheduler
initial_lr: 0.00006 # for effective batch_size=8
multi_gpu: DDP
num_gpus: 4
---
Src/Super Gradients/Recipes/Cityscapes Segformer B2.Yaml (src/super_gradients/recipes/cityscapes_segformer_b2.yaml)
defaults:
- cityscapes_segformer
- _self_
- variable_setup
architecture: segformer_b2
checkpoint_params:
checkpoint_path:
load_backbone: True
load_weights_only: True
strict_load: no_key_matching
load_checkpoint: False
experiment_name: ${architecture}_cityscapes
training_hyperparams:
max_epochs: 2
lr_mode: PolyLRScheduler
initial_lr: 0.00006 # for effective batch_size=8
multi_gpu: DDP
num_gpus: 4
---
Src/Super Gradients/Recipes/Cityscapes Segformer B3.Yaml (src/super_gradients/recipes/cityscapes_segformer_b3.yaml)
defaults:
- cityscapes_segformer
- _self_
- variable_setup
architecture: segformer_b3
checkpoint_params:
checkpoint_path:
load_backbone: True
load_weights_only: True
strict_load: no_key_matching
load_checkpoint: False
experiment_name: ${architecture}_cityscapes
training_hyperparams:
max_epochs: 2
lr_mode: PolyLRScheduler
initial_lr: 0.00006 # for effective batch_size=8
multi_gpu: DDP
num_gpus: 4
---
Src/Super Gradients/Recipes/Cityscapes Segformer B4.Yaml (src/super_gradients/recipes/cityscapes_segformer_b4.yaml)
defaults:
- cityscapes_segformer
- _self_
- variable_setup
architecture: segformer_b4
checkpoint_params:
checkpoint_path:
load_backbone: True
load_weights_only: True
strict_load: no_key_matching
load_checkpoint: False
experiment_name: ${architecture}_cityscapes
training_hyperparams:
max_epochs: 2
lr_mode: PolyLRScheduler
initial_lr: 0.00006 # for effective batch_size=8
mixed_precision: True
multi_gpu: DDP
num_gpus: 4
---
Src/Super Gradients/Recipes/Cityscapes Segformer B5.Yaml (src/super_gradients/recipes/cityscapes_segformer_b5.yaml)
defaults:
- cityscapes_segformer
- _self_
- variable_setup
architecture: segformer_b5
checkpoint_params:
checkpoint_path:
load_backbone: True
load_weights_only: True
strict_load: no_key_matching
load_checkpoint: False
experiment_name: ${architecture}_cityscapes
training_hyperparams:
max_epochs: 2
lr_mode: PolyLRScheduler
initial_lr: 0.00006 # for effective batch_size=8
mixed_precision: True
multi_gpu: DDP
num_gpus: 4
---
Src/Super Gradients/Recipes/Cityscapes Stdc Base.Yaml (src/super_gradients/recipes/cityscapes_stdc_base.yaml)
STDC Base training params
defaults:
- training_hyperparams: cityscapes_default_train_params
- dataset_params: cityscapes_dataset_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
train_dataloader: cityscapes_train
val_dataloader: cityscapes_val
data_loader_num_workers: 10
arch_params:
num_classes: 19
use_aux_heads: True
checkpoint_params:
checkpoint_path:
load_backbone: True
load_weights_only: True
strict_load: no_key_matching
architecture: stdc1_seg
experiment_name: ${architecture}_cityscapes
training_hyperparams:
sync_bn: True
multi_gpu: DDP
---
Src/Super Gradients/Recipes/Cityscapes Stdc Seg50.Yaml (src/super_gradients/recipes/cityscapes_stdc_seg50.yaml)
STDC segmentation training example with Cityscapes dataset.
Reproduction and refinement of paper: Rethinking BiSeNet For Real-time Semantic Segmentation.
#Instructions:
0. Make sure that the data is stored in dataset_params.dataset_dir or add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below (feel free to check ReadMe)
1. Move to the project root (where you will find the ReadMe and src folder)
2. Run the command:
STDC1-Seg50: python -m super_gradients.train_from_recipe --config-name=cityscapes_stdc_seg50
STDC2-Seg50: python -m super_gradients.train_from_recipe --config-name=cityscapes_stdc_seg50 architecture=stdc2_seg
Note: add "checkpoint_params.checkpoint_path=<stdc1-backbone-pretrained-path>" to use pretrained backbone
#
#
#
Validation mIoU - Cityscapes, training time:
STDC1-Seg50: input-size: [512, 1024] mIoU: 75.11 2 X RTX A5000, 20 H
STDC2-Seg50: input-size: [512, 1024] mIoU: 76.44 2 X RTX A5000, 23 H
#
Official git repo:
https://github.com/MichaelFan01/STDC-Seg
Paper:
https://arxiv.org/abs/2104.13188
#
Pretrained checkpoints:
Backbones- downloaded from the author's official repo.
https://deci-pretrained-models.s3.amazonaws.com/stdc_backbones/stdc1_imagenet_pretrained.pth
https://deci-pretrained-models.s3.amazonaws.com/stdc_backbones/stdc2_imagenet_pretrained.pth
#
Logs, tensorboards and network checkpoints:
STDC1-Seg50: https://deci-pretrained-models.s3.amazonaws.com/cityscapes_stdc1_seg50_dice_edge/
STDC2-Seg50: https://deci-pretrained-models.s3.amazonaws.com/cityscapes_stdc2_seg50_dice_edge/
#
Learning rate and batch size parameters, using 2 RTX A5000 with DDP:
STDC1-Seg50: input-size: [512, 1024] initial_lr: 0.01 batch-size: 16 * 2gpus = 32
STDC2-Seg50: input-size: [512, 1024] initial_lr: 0.01 batch-size: 16 * 2gpus = 32
#
Comments:
* Pretrained backbones were used.
defaults:
- training_hyperparams: cityscapes_default_train_params
- dataset_params: cityscapes_stdc_seg50_dataset_params # TODO: uncomment after DatasetInterface refactor
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
train_dataloader: cityscapes_train
val_dataloader: cityscapes_val
architecture: stdc1_seg
arch_params:
num_classes: 19
use_aux_heads: True
checkpoint_params:
checkpoint_path:
load_backbone: True
load_weights_only: True
strict_load: no_key_matching
training_hyperparams:
initial_lr:
cp: 0.01
default: 0.1
sync_bn: True
loss:
DiceCEEdgeLoss:
num_classes: 19
ignore_index: 19
weights: [ 1., 0.6, 0.4, 1. ]
dice_ce_weights: [ 1., 1. ]
ce_edge_weights: [ .5, .5 ]
edge_kernel: 3
multi_gpu: DDP
num_gpus: 2
experiment_name: ${architecture}50_cityscapes
---
Src/Super Gradients/Recipes/Cityscapes Stdc Seg75.Yaml (src/super_gradients/recipes/cityscapes_stdc_seg75.yaml)
STDC segmentation training example with Cityscapes dataset.
Reproduction and refinement of paper: Rethinking BiSeNet For Real-time Semantic Segmentation.
#Instructions:
0. Make sure that the data is stored in dataset_params.dataset_dir or add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below (feel free to check ReadMe)
1. Move to the project root (where you will find the ReadMe and src folder)
2. Run the command:
STDC1-Seg75: python -m super_gradients.train_from_recipe --config-name=cityscapes_stdc_seg75
STDC2-Seg75: python -m super_gradients.train_from_recipe --config-name=cityscapes_stdc_seg75 architecture=stdc2_seg
Note: add "external_checkpoint_path=<stdc1-backbone-pretrained-path>" to use pretrained backbone
#
#
#
Validation mIoU - Cityscapes, training time:
STDC1-Seg75: input-size: [768, 1536] mIoU: 76.87 4 X RTX A5000, 29 H, early stopped after 711 epochs
STDC2-Seg75: input-size: [768, 1536] mIoU: 78.93 2 X RTX A5000, 29 H, early stopped after 530 epochs
#
Official git repo:
https://github.com/MichaelFan01/STDC-Seg
Paper:
https://arxiv.org/abs/2104.13188
#
Pretrained checkpoints:
Backbones- downloaded from the author's official repo.
https://deci-pretrained-models.s3.amazonaws.com/stdc_backbones/stdc1_imagenet_pretrained.pth
https://deci-pretrained-models.s3.amazonaws.com/stdc_backbones/stdc2_imagenet_pretrained.pth
#
Logs, tensorboards and network checkpoints:
https://deci-pretrained-models.s3.amazonaws.com/stdc1_seg75_cityscapes/
https://deci-pretrained-models.s3.amazonaws.com/stdc2_seg75_cityscapes/
#
#
Learning rate and batch size parameters, using 4 GeForce RTX 2080 Ti with DDP:
STDC1-Seg75: input-size: [768, 1536] initial_lr: 0.005 batch-size: 4 * 4gpus = 16
STDC2-Seg75: input-size: [768, 1536] initial_lr: 0.005 batch-size: 8 * 2gpus = 16
#
Comments:
* Pretrained backbones were used.
* Results with Deci code are higher than original implementation, mostly thanks to changes in Detail loss and
module, different auxiliary feature maps and different loss weights.
defaults:
- training_hyperparams: cityscapes_default_train_params
- dataset_params: cityscapes_stdc_seg75_dataset_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
train_dataloader: cityscapes_train
val_dataloader: cityscapes_val
architecture: stdc1_seg
arch_params:
num_classes: 19
use_aux_heads: True
checkpoint_params:
checkpoint_path:
load_backbone: True
load_weights_only: True
strict_load: no_key_matching
training_hyperparams:
initial_lr:
cp: 0.005
default: 0.05
sync_bn: True
loss:
STDCLoss:
num_classes: 19
ignore_index: 19
mining_percent: 0.0625 # mining percentage is 1/16 of pixels following original implementation.
weights: [ 1., 0.6, 0.4, 1. ]
multi_gpu: DDP
num_gpus: 4
experiment_name: ${architecture}75_cityscapes
---
Src/Super Gradients/Recipes/Coco2017 Pose Dekr Rescoring.Yaml (src/super_gradients/recipes/coco2017_pose_dekr_rescoring.yaml)
This file contains the recipe to train rescoring model for DERK pose estimation model.
See documentation/source/PoseEstimation.md#Rescoring for a detailed explanation of the rescoring model.
#
Important:
If you want to train your own rescoring model, you need to generate the rescoring data first:
python -m super_gradients.scripts.generate_rescoring_training_data --config-name=script_generate_rescoring_data_dekr_coco2017 rescoring_data_dir=OUTPUT_DATA_DIR checkpoint=PATH_TO_TRAINED_MODEL_CHECKPOINT
#
Usage:
python -m super_gradients.train_from_recipe --config-name coco2017_pose_dekr_rescoring \
dataset_params.train_dataset_params.pkl_file=OUTPUT_DATA_DIR/rescoring_data_train.pkl \
dataset_params.val_dataset_params.pkl_file=OUTPUT_DATA_DIR/rescoring_data_valid.pkl
defaults:
- training_hyperparams: coco2017_rescoring_train_params
- dataset_params: coco_pose_estimation_rescoring_dataset_params
- arch_params: pose_dekr_coco_rescoring_arch_params
- checkpoint_params: default_checkpoint_params
- _self_
resume: False
architecture: pose_rescoring_coco
multi_gpu: Off
num_gpus: 1
experiment_suffix: ""
experiment_name: coco2017_pose_dekr_rescoring_${architecture}_${experiment_suffix}
ckpt_root_dir:
train_dataloader: coco2017_rescoring_train
val_dataloader: coco2017_rescoring_val
training_hyperparams:
resume: ${resume}
THE FOLLOWING PARAMS ARE DIRECTLY USED BY HYDRA
hydra:
run:
# Set the output directory (i.e. where .hydra folder that logs all the input params will be generated)
dir: ${hydra_output_dir:${ckpt_root_dir}, ${experiment_name}}
---
Src/Super Gradients/Recipes/Coco2017 Pose Dekr W32 No Dc.Yaml (src/super_gradients/recipes/coco2017_pose_dekr_w32_no_dc.yaml)
DEKR training example with COCO dataset.
Reproduction and refinement of paper: Bottom-Up Human Pose Estimation Via Disentangled Keypoint Regression.
#
Note: Original DEKR architecture using deformable convolutions. This recipe uses standard convolutions to enable
model be exportable to ONNX.
#
Recipe runs with batch size = 24 X 8 gpus = 192.
#
Instructions:
0. Make sure that the data is stored in dataset_params.dataset_dir or add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below (feel free to check ReadMe)
1. Move to the project root (where you will find the ReadMe and src folder)
2. Make sure you've downloaded pretrained backbone weights from https://1drv.ms/u/s!Aus8VCZ_C_33dYBMemi9xOUFR0w to project root (See line 55).
3. Run the command:
DEKR-W32-NO-DC: python -m super_gradients.train_from_recipe --config-name=coco2017_pose_dekr_w32_no_dc checkpoint_params.checkpoint_path=hrnetv2_w32_imagenet_pretrained.pth
#
#
Validation AP (Without flip augmentation and rescoring) - COCO, training time:
DEKR-W32-NO-DC: input-size: [640, 640] AP: 63.08 (Regular training) 8 X RTX A5000 - 21h
#
Scores with flip TTA and rescoring (Using best model from above):
DEKR-W32-NO-DC: input-size: [640, 640] AP: 64.96 (With Flip TTA)
DEKR-W32-NO-DC: input-size: [640, 640] AP: 67.34 (With Flip TTA and Rescoring)
#
Rescoring:
See coco2017_pose_dekr_rescoring.yaml recipe and documentation/source/PoseEstimation.md#Rescoring section of the documentation.
#
Official git repo:
https://github.com/HRNet/DEKR
Paper:
https://arxiv.org/abs/2104.02300
#
#
Comments:
* Pretrained backbones were used.
* In DEKR-W32-NO-DC A suffix "NO-DC" stands for "No deformable convolutions".
defaults:
- training_hyperparams: coco2017_dekr_pose_train_params
- dataset_params: coco_pose_estimation_dekr_dataset_params
- arch_params: pose_dekr_w32_no_dc_arch_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
architecture: dekr_w32_no_dc
multi_gpu: DDP
num_gpus: 8
experiment_suffix: ""
experiment_name: coco2017_pose_${architecture}${experiment_suffix}
ckpt_root_dir:
train_dataloader: coco2017_pose_train
val_dataloader: coco2017_pose_val
arch_params:
num_classes: ${dataset_params.num_joints}
checkpoint_params:
# Original training recipe uses pretrained weights for HRNet on ImageNet.
# You will need to download the pretrained weights from the original repo and place
# them in external_checkpoint_path param.
# Download weights from this url https://1drv.ms/u/s!Aus8VCZ_C_33dYBMemi9xOUFR0w
checkpoint_path: # <location of the downloaded hrnetv2_w32_imagenet_pretrained.pth>
strict_load:
_target_: super_gradients.training.sg_trainer.StrictLoad
value: key_matching
dataset_params:
train_dataloader_params:
batch_size: 24
val_dataloader_params:
batch_size: 32
---
Src/Super Gradients/Recipes/Coco2017 Ppyoloe L.Yaml (src/super_gradients/recipes/coco2017_ppyoloe_l.yaml)
PP-Yolo-E Detection training on COCO2017 Dataset:
PP-Yolo-E trained in 640x640
Checkpoints + tensorboards: https://deci-pretrained-models.s3.amazonaws.com/ppyoloe_coco/
Recipe runs with batch size = 20 X 8 gpus = 160.
Instructions:
0. Make sure that the data is stored in dataset_params.dataset_dir or add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below (feel free to check ReadMe)
1. Move to the project root (where you will find the ReadMe and src folder)
2. Run the command you want:
ppyoloe_s: python -m super_gradients.train_from_recipe --config-name=coco2017_ppyoloe_s
ppyoloe_m: python -m super_gradients.train_from_recipe --config-name=coco2017_ppyoloe_m
ppyoloe_l: python -m super_gradients.train_from_recipe --config-name=coco2017_ppyoloe_l
ppyoloe_x: python -m super_gradients.train_from_recipe --config-name=coco2017_ppyoloe_x
#
Training times and accuracies ([email protected] (COCO API, confidence 0.001, IoU threshold 0.6, test on 640x640 images):
ppyoloe_s: 37h on 8 NVIDIA GeForce RTX 3090, mAP: 42.52 (val)
ppyoloe_m: 58h on 8 NVIDIA GeForce RTX 3090, mAP: 47.11 (val)
ppyoloe_l: 115h on 8 NVIDIA GeForce RTX 3090, mAP: 49.48 (val)
ppyoloe_x: 240h on 8 NVIDIA GeForce RTX 3090, mAP: 51.15 (val)
#
defaults:
- training_hyperparams: coco2017_ppyoloe_train_params
- dataset_params: coco_detection_ppyoloe_dataset_params
- arch_params: ppyoloe_l_arch_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
train_dataloader: coco2017_train_ppyoloe
val_dataloader: coco2017_val_ppyoloe
load_checkpoint: False
resume: False
dataset_params:
train_dataloader_params:
batch_size: 20
training_hyperparams:
resume: ${resume}
mixed_precision: True
initial_lr: 1e-3
architecture: pp_yoloe_l
multi_gpu: DDP
num_gpus: 8
experiment_suffix: ""
experiment_name: coco2017_${architecture}${experiment_suffix}
---
Src/Super Gradients/Recipes/Coco2017 Ppyoloe M.Yaml (src/super_gradients/recipes/coco2017_ppyoloe_m.yaml)
PP-Yolo-E Detection training on COCO2017 Dataset:
PP-Yolo-E trained in 640x640
Checkpoints + tensorboards: https://deci-pretrained-models.s3.amazonaws.com/ppyoloe_coco/
Recipe runs with batch size = 24 X 8 gpus = 192.
Instructions:
0. Make sure that the data is stored in dataset_params.dataset_dir or add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below (feel free to check ReadMe)
1. Move to the project root (where you will find the ReadMe and src folder)
2. Run the command you want:
ppyoloe_s: python -m super_gradients.train_from_recipe --config-name=coco2017_ppyoloe_s
ppyoloe_m: python -m super_gradients.train_from_recipe --config-name=coco2017_ppyoloe_m
ppyoloe_l: python -m super_gradients.train_from_recipe --config-name=coco2017_ppyoloe_l
ppyoloe_x: python -m super_gradients.train_from_recipe --config-name=coco2017_ppyoloe_x
#
Training times and accuracies ([email protected] (COCO API, confidence 0.001, IoU threshold 0.6, test on 640x640 images):
ppyoloe_s: 37h on 8 NVIDIA GeForce RTX 3090, mAP: 42.52 (val)
ppyoloe_m: 58h on 8 NVIDIA GeForce RTX 3090, mAP: 47.11 (val)
ppyoloe_l: 115h on 8 NVIDIA GeForce RTX 3090, mAP: 49.48 (val)
ppyoloe_x: 240h on 8 NVIDIA GeForce RTX 3090, mAP: 51.15 (val)
#
defaults:
- training_hyperparams: coco2017_ppyoloe_train_params
- dataset_params: coco_detection_ppyoloe_dataset_params
- arch_params: ppyoloe_m_arch_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
train_dataloader: coco2017_train_ppyoloe
val_dataloader: coco2017_val_ppyoloe
load_checkpoint: False
resume: False
dataset_params:
train_dataloader_params:
batch_size: 24
training_hyperparams:
resume: ${resume}
mixed_precision: True
initial_lr: 1e-3
architecture: pp_yoloe_m
multi_gpu: DDP
num_gpus: 8
experiment_suffix: ""
experiment_name: coco2017_${architecture}${experiment_suffix}
---
Src/Super Gradients/Recipes/Coco2017 Ppyoloe S.Yaml (src/super_gradients/recipes/coco2017_ppyoloe_s.yaml)
PP-Yolo-E Detection training on COCO2017 Dataset:
PP-Yolo-E trained in 640x640
Recipe runs with batch size = 32 X 8 gpus = 256.
Instructions:
0. Make sure that the data is stored in dataset_params.dataset_dir or add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below (feel free to check ReadMe)
1. Move to the project root (where you will find the ReadMe and src folder)
2. Run the command you want:
ppyoloe_s: python -m super_gradients.train_from_recipe --config-name=coco2017_ppyoloe_s
ppyoloe_m: python -m super_gradients.train_from_recipe --config-name=coco2017_ppyoloe_m
ppyoloe_l: python -m super_gradients.train_from_recipe --config-name=coco2017_ppyoloe_l
ppyoloe_x: python -m super_gradients.train_from_recipe --config-name=coco2017_ppyoloe_x
#
Training times and accuracies ([email protected] (COCO API, confidence 0.001, IoU threshold 0.6, test on 640x640 images):
ppyoloe_s: 37h on 8 NVIDIA GeForce RTX 3090, mAP: 42.52 (val)
ppyoloe_m: 58h on 8 NVIDIA GeForce RTX 3090, mAP: 47.11 (val)
ppyoloe_l: 115h on 8 NVIDIA GeForce RTX 3090, mAP: 49.48 (val)
ppyoloe_x: 240h on 8 NVIDIA GeForce RTX 3090, mAP: 51.15 (val)
#
defaults:
- training_hyperparams: coco2017_ppyoloe_train_params
- dataset_params: coco_detection_ppyoloe_dataset_params
- arch_params: ppyoloe_s_arch_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
train_dataloader: coco2017_train_ppyoloe
val_dataloader: coco2017_val_ppyoloe
load_checkpoint: False
resume: False
dataset_params:
train_dataloader_params:
batch_size: 32
training_hyperparams:
resume: ${resume}
mixed_precision: True
architecture: pp_yoloe_s
multi_gpu: DDP
num_gpus: 8
experiment_suffix: ""
experiment_name: coco2017_${architecture}${experiment_suffix}
---
Src/Super Gradients/Recipes/Coco2017 Ppyoloe X.Yaml (src/super_gradients/recipes/coco2017_ppyoloe_x.yaml)
PP-Yolo-E Detection training on COCO2017 Dataset:
PP-Yolo-E trained in 640x640
Checkpoints + tensorboards: https://deci-pretrained-models.s3.amazonaws.com/ppyoloe_coco/
Recipe runs with batch size = 16 X 8 gpus = 128.
Instructions:
0. Make sure that the data is stored in dataset_params.dataset_dir or add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below (feel free to check ReadMe)
1. Move to the project root (where you will find the ReadMe and src folder)
2. Run the command you want:
ppyoloe_s: python -m super_gradients.train_from_recipe --config-name=coco2017_ppyoloe_s
ppyoloe_m: python -m super_gradients.train_from_recipe --config-name=coco2017_ppyoloe_m
ppyoloe_l: python -m super_gradients.train_from_recipe --config-name=coco2017_ppyoloe_l
ppyoloe_x: python -m super_gradients.train_from_recipe --config-name=coco2017_ppyoloe_x
#
Training times and accuracies ([email protected] (COCO API, confidence 0.001, IoU threshold 0.6, test on 640x640 images):
ppyoloe_s: 37h on 8 NVIDIA GeForce RTX 3090, mAP: 42.52 (val)
ppyoloe_m: 58h on 8 NVIDIA GeForce RTX 3090, mAP: 47.11 (val)
ppyoloe_l: 115h on 8 NVIDIA GeForce RTX 3090, mAP: 49.48 (val)
ppyoloe_x: 240h on 8 NVIDIA GeForce RTX 3090, mAP: 51.15 (val)
#
defaults:
- training_hyperparams: coco2017_ppyoloe_train_params
- dataset_params: coco_detection_ppyoloe_dataset_params
- arch_params: ppyoloe_x_arch_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
train_dataloader: coco2017_train_ppyoloe
val_dataloader: coco2017_val_ppyoloe
load_checkpoint: False
resume: False
dataset_params:
train_dataloader_params:
batch_size: 16
training_hyperparams:
resume: ${resume}
mixed_precision: True
architecture: pp_yoloe_x
multi_gpu: DDP
num_gpus: 8
experiment_suffix: ""
experiment_name: coco2017_${architecture}${experiment_suffix}
---
Src/Super Gradients/Recipes/Coco2017 Ssd Lite Mobilenet V2.Yaml (src/super_gradients/recipes/coco2017_ssd_lite_mobilenet_v2.yaml)
SSD MobileNetV2 Detection training on CoCo2017 Dataset:
Trained in 320x320 [email protected]@0.95 (COCO API, confidence 0.001, IoU threshold 0.6, test on 320x320 images) ~20.41
Checkpoint path: https://deci-pretrained-models.s3.amazonaws.com/ssd_lite_mobilenet_v2/coco2017/2022-11-28/average_model.pth
(trained with stride_16_plus_big)
Hardware: 4 NVIDIA RTX 2080Ti
Training time: Β±35 hours
#
Instructions:
0. Make sure that the data is stored in dataset_params.dataset_dir or add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below (feel free to check ReadMe)
1. Move to the project root (where you will find the ReadMe and src folder)
2. Run the command:
python -m super_gradients.train_from_recipe --config-name=coco2017_ssd_lite_mobilenet_v2
NOTE:
Anchors will be selected based on validation resolution and anchors_name
To switch between anchors, set anchors_name to something else defined in the anchors section
e.g.
python -m super_gradients.train_from_recipe --config-name=coco2017_ssd_lite_mobilenet_v2 anchors_name=stride_16_plus
defaults:
- training_hyperparams: coco2017_ssd_lite_mobilenet_v2_train_params
- dataset_params: coco_detection_ssd_lite_mobilenet_v2_dataset_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
- anchors: ssd_anchors
train_dataloader: coco2017_train
val_dataloader: coco2017_val
architecture: ssd_lite_mobilenet_v2
data_loader_num_workers: 8
experiment_suffix: res320
experiment_name: ${architecture}_coco_${experiment_suffix}
anchors_resolution: 320x320
anchors_name: stride_16_plus_big
dboxes: ${anchors.${anchors_resolution}.${anchors_name}}
arch_params:
num_classes: 80
anchors: ${dboxes}
resume: False
training_hyperparams:
resume: ${resume}
criterion_params:
alpha: 1.0
dboxes: ${dboxes}
multi_gpu: DDP
num_gpus: 4
---
Src/Super Gradients/Recipes/Coco2017 Yolo Nas Pose L.Yaml (src/super_gradients/recipes/coco2017_yolo_nas_pose_l.yaml)
YoloNASPose training on COCO2017 Dataset
All YoloNASPose models trained in 640x640 resolution
#
Instructions:
0. Have super-gradients installed (pip install super-gradients==3.3 or clone the repo and pip install -e .)
1. Make sure that the data is stored folder specified at dataset_params.dataset_dir (Default is /data/coco) or
add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below
2. Run the command to start the training:
yolo_nas_pose_n: python -m super_gradients.train_from_recipe --config-name=coco2017_yolo_nas_pose_n
yolo_nas_pose_s: python -m super_gradients.train_from_recipe --config-name=coco2017_yolo_nas_pose_s
yolo_nas_pose_m: python -m super_gradients.train_from_recipe --config-name=coco2017_yolo_nas_pose_m
yolo_nas_pose_l: python -m super_gradients.train_from_recipe --config-name=coco2017_yolo_nas_pose_l
#
Training times and accuracies ([email protected] (COCO API, confidence 0.01, IoU threshold 0.7, test on original resolution):
yolo_nas_pose_n: 93h on 8 NVIDIA GeForce RTX 3090, AP: 59.68 (val)
yolo_nas_pose_s: 52h on 8 NVIDIA GeForce RTX 3090, AP: 64.15 (val)
yolo_nas_pose_m: 57h on 8 NVIDIA GeForce RTX 3090, AP: 67.87 (val)
yolo_nas_pose_l: 80h on 8 NVIDIA GeForce RTX 3090, AP: 68.24 (val)
#
Offline evaluation using COCOEval for L variant:
Average Precision (AP) @[ IoU=0.50:0.95 | area= all | maxDets= 20 ] = 0.682
Average Precision (AP) @[ IoU=0.50 | area= all | maxDets= 20 ] = 0.891
Average Precision (AP) @[ IoU=0.75 | area= all | maxDets= 20 ] = 0.752
Average Precision (AP) @[ IoU=0.50:0.95 | area=medium | maxDets= 20 ] = 0.631
Average Precision (AP) @[ IoU=0.50:0.95 | area= large | maxDets= 20 ] = 0.766
Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets= 20 ] = 0.735
Average Recall (AR) @[ IoU=0.50 | area= all | maxDets= 20 ] = 0.924
Average Recall (AR) @[ IoU=0.75 | area= all | maxDets= 20 ] = 0.799
Average Recall (AR) @[ IoU=0.50:0.95 | area=medium | maxDets= 20 ] = 0.683
Average Recall (AR) @[ IoU=0.50:0.95 | area= large | maxDets= 20 ] = 0.811
defaults:
- training_hyperparams: coco2017_yolo_nas_pose_train_params
- dataset_params: coco_pose_estimation_yolo_nas_mosaic_heavy_dataset_params
- arch_params: yolo_nas_pose_l_arch_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
architecture: yolo_nas_pose_l
multi_gpu: DDP
num_gpus: 8
experiment_suffix: ""
experiment_name: coco2017_${architecture}_${experiment_suffix}_${dataset_params.dataset_params_suffix}
arch_params:
num_classes: ${dataset_params.num_joints}
training_hyperparams:
initial_lr: 8e-5
criterion_params:
dfl_loss_weight: 0.5
pose_reg_loss_weight: 10
ema_params:
decay_type: exp
beta: 50
checkpoint_params:
# For training Yolo-NAS-S pose estimation model we use pretrained weights for Yolo-NAS-S object detection model.
# By setting strict_load: key_matching we load only those weights that match the keys of the model.
checkpoint_path: https://sghub.deci.ai/models/yolo_nas_l_coco.pth
strict_load:
_target_: super_gradients.training.sg_trainer.StrictLoad
value: key_matching
dataset_params:
mosaic_prob: 0.5
train_dataloader_params:
batch_size: 24
val_dataloader_params:
batch_size: 24
---
Src/Super Gradients/Recipes/Coco2017 Yolo Nas Pose M.Yaml (src/super_gradients/recipes/coco2017_yolo_nas_pose_m.yaml)
YoloNASPose training on COCO2017 Dataset
All YoloNASPose models trained in 640x640 resolution
#
Instructions:
0. Have super-gradients installed (pip install super-gradients==3.3 or clone the repo and pip install -e .)
1. Make sure that the data is stored folder specified at dataset_params.dataset_dir (Default is /data/coco) or
add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below
2. Run the command to start the training:
yolo_nas_pose_n: python -m super_gradients.train_from_recipe --config-name=coco2017_yolo_nas_pose_n
yolo_nas_pose_s: python -m super_gradients.train_from_recipe --config-name=coco2017_yolo_nas_pose_s
yolo_nas_pose_m: python -m super_gradients.train_from_recipe --config-name=coco2017_yolo_nas_pose_m
yolo_nas_pose_l: python -m super_gradients.train_from_recipe --config-name=coco2017_yolo_nas_pose_l
#
Training times and accuracies ([email protected] (COCO API, confidence 0.01, IoU threshold 0.7, test on original resolution):
yolo_nas_pose_n: 93h on 8 NVIDIA GeForce RTX 3090, AP: 59.68 (val)
yolo_nas_pose_s: 52h on 8 NVIDIA GeForce RTX 3090, AP: 64.15 (val)
yolo_nas_pose_m: 57h on 8 NVIDIA GeForce RTX 3090, AP: 67.87 (val)
yolo_nas_pose_l: 80h on 8 NVIDIA GeForce RTX 3090, AP: 68.24 (val)
#
Offline evaluation using COCOEval for M variant:
Average Precision (AP) @[ IoU=0.50:0.95 | area= all | maxDets= 20 ] = 0.679
Average Precision (AP) @[ IoU=0.50 | area= all | maxDets= 20 ] = 0.887
Average Precision (AP) @[ IoU=0.75 | area= all | maxDets= 20 ] = 0.745
Average Precision (AP) @[ IoU=0.50:0.95 | area=medium | maxDets= 20 ] = 0.622
Average Precision (AP) @[ IoU=0.50:0.95 | area= large | maxDets= 20 ] = 0.771
Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets= 20 ] = 0.730
Average Recall (AR) @[ IoU=0.50 | area= all | maxDets= 20 ] = 0.919
Average Recall (AR) @[ IoU=0.75 | area= all | maxDets= 20 ] = 0.792
Average Recall (AR) @[ IoU=0.50:0.95 | area=medium | maxDets= 20 ] = 0.672
Average Recall (AR) @[ IoU=0.50:0.95 | area= large | maxDets= 20 ] = 0.813
defaults:
- training_hyperparams: coco2017_yolo_nas_pose_train_params
- dataset_params: coco_pose_estimation_yolo_nas_mosaic_heavy_dataset_params
- arch_params: yolo_nas_pose_m_arch_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
architecture: yolo_nas_pose_m
multi_gpu: DDP
num_gpus: 8
experiment_suffix: ""
experiment_name: coco2017_${architecture}_${experiment_suffix}_${dataset_params.dataset_params_suffix}
arch_params:
num_classes: ${dataset_params.num_joints}
training_hyperparams:
initial_lr: 1e-4 # Scale factor to account reduced batch size
criterion_params:
dfl_loss_weight: 0.5
pose_reg_loss_weight: 10
ema_params:
decay_type: exp
beta: 50
checkpoint_params:
# For training Yolo-NAS-S pose estimation model we use pretrained weights for Yolo-NAS-S object detection model.
# By setting strict_load: key_matching we load only those weights that match the keys of the model.
checkpoint_path: https://sghub.deci.ai/models/yolo_nas_m_coco.pth
strict_load:
_target_: super_gradients.training.sg_trainer.StrictLoad
value: key_matching
dataset_params:
mosaic_prob: 0.5
train_dataloader_params:
batch_size: 32
val_dataloader_params:
batch_size: 32
---
Src/Super Gradients/Recipes/Coco2017 Yolo Nas Pose N.Yaml (src/super_gradients/recipes/coco2017_yolo_nas_pose_n.yaml)
YoloNASPose training on COCO2017 Dataset
All YoloNASPose models trained in 640x640 resolution
#
Instructions:
0. Have super-gradients installed (pip install super-gradients==3.3 or clone the repo and pip install -e .)
1. Make sure that the data is stored folder specified at dataset_params.dataset_dir (Default is /data/coco) or
add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below
2. Run the command to start the training:
yolo_nas_pose_n: python -m super_gradients.train_from_recipe --config-name=coco2017_yolo_nas_pose_n
yolo_nas_pose_s: python -m super_gradients.train_from_recipe --config-name=coco2017_yolo_nas_pose_s
yolo_nas_pose_m: python -m super_gradients.train_from_recipe --config-name=coco2017_yolo_nas_pose_m
yolo_nas_pose_l: python -m super_gradients.train_from_recipe --config-name=coco2017_yolo_nas_pose_l
#
Training times and accuracies ([email protected] (COCO API, confidence 0.01, IoU threshold 0.7, test on original resolution):
yolo_nas_pose_n: 93h on 8 NVIDIA GeForce RTX 3090, AP: 59.68 (val)
yolo_nas_pose_s: 52h on 8 NVIDIA GeForce RTX 3090, AP: 64.15 (val)
yolo_nas_pose_m: 57h on 8 NVIDIA GeForce RTX 3090, AP: 67.87 (val)
yolo_nas_pose_l: 80h on 8 NVIDIA GeForce RTX 3090, AP: 68.24 (val)
#
Offline evaluation using COCOEval for N variant:
Average Precision (AP) @[ IoU=0.50:0.95 | area= all | maxDets= 20 ] = 0.597
Average Precision (AP) @[ IoU=0.50 | area= all | maxDets= 20 ] = 0.832
Average Precision (AP) @[ IoU=0.75 | area= all | maxDets= 20 ] = 0.657
Average Precision (AP) @[ IoU=0.50:0.95 | area=medium | maxDets= 20 ] = 0.540
Average Precision (AP) @[ IoU=0.50:0.95 | area= large | maxDets= 20 ] = 0.685
Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets= 20 ] = 0.656
Average Recall (AR) @[ IoU=0.50 | area= all | maxDets= 20 ] = 0.879
Average Recall (AR) @[ IoU=0.75 | area= all | maxDets= 20 ] = 0.713
Average Recall (AR) @[ IoU=0.50:0.95 | area=medium | maxDets= 20 ] = 0.598
Average Recall (AR) @[ IoU=0.50:0.95 | area= large | maxDets= 20 ] = 0.738
defaults:
- training_hyperparams: coco2017_yolo_nas_pose_train_params
- dataset_params: coco_pose_estimation_yolo_nas_mosaic_dataset_params
- arch_params: yolo_nas_pose_n_arch_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
architecture: yolo_nas_pose_n
multi_gpu: DDP
num_gpus: 8
experiment_suffix: ""
experiment_name: coco2017_${architecture}_${experiment_suffix}_${dataset_params.dataset_params_suffix}
arch_params:
num_classes: ${dataset_params.num_joints}
dataset_params:
mosaic_prob: 0.5
train_dataloader_params:
batch_size: 60
val_dataloader_params:
batch_size: 60
---
Src/Super Gradients/Recipes/Coco2017 Yolo Nas Pose S.Yaml (src/super_gradients/recipes/coco2017_yolo_nas_pose_s.yaml)
YoloNASPose training on COCO2017 Dataset
All YoloNASPose models trained in 640x640 resolution
#
Instructions:
0. Have super-gradients installed (pip install super-gradients==3.3 or clone the repo and pip install -e .)
1. Make sure that the data is stored folder specified at dataset_params.dataset_dir (Default is /data/coco) or
add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below
2. Run the command to start the training:
yolo_nas_pose_n: python -m super_gradients.train_from_recipe --config-name=coco2017_yolo_nas_pose_n
yolo_nas_pose_s: python -m super_gradients.train_from_recipe --config-name=coco2017_yolo_nas_pose_s
yolo_nas_pose_m: python -m super_gradients.train_from_recipe --config-name=coco2017_yolo_nas_pose_m
yolo_nas_pose_l: python -m super_gradients.train_from_recipe --config-name=coco2017_yolo_nas_pose_l
#
Training times and accuracies ([email protected] (COCO API, confidence 0.01, IoU threshold 0.7, test on original resolution):
yolo_nas_pose_n: 93h on 8 NVIDIA GeForce RTX 3090, AP: 59.68 (val)
yolo_nas_pose_s: 52h on 8 NVIDIA GeForce RTX 3090, AP: 64.15 (val)
yolo_nas_pose_m: 57h on 8 NVIDIA GeForce RTX 3090, AP: 67.87 (val)
yolo_nas_pose_l: 80h on 8 NVIDIA GeForce RTX 3090, AP: 68.24 (val)
#
Offline evaluation using COCOEval for S variant:
Average Precision (AP) @[ IoU=0.50:0.95 | area= all | maxDets= 20 ] = 0.642
Average Precision (AP) @[ IoU=0.50 | area= all | maxDets= 20 ] = 0.856
Average Precision (AP) @[ IoU=0.75 | area= all | maxDets= 20 ] = 0.703
Average Precision (AP) @[ IoU=0.50:0.95 | area=medium | maxDets= 20 ] = 0.594
Average Precision (AP) @[ IoU=0.50:0.95 | area= large | maxDets= 20 ] = 0.723
Average Recall (AR) @[ IoU=0.50:0.95 | area= all | maxDets= 20 ] = 0.702
Average Recall (AR) @[ IoU=0.50 | area= all | maxDets= 20 ] = 0.901
Average Recall (AR) @[ IoU=0.75 | area= all | maxDets= 20 ] = 0.759
Average Recall (AR) @[ IoU=0.50:0.95 | area=medium | maxDets= 20 ] = 0.650
Average Recall (AR) @[ IoU=0.50:0.95 | area= large | maxDets= 20 ] = 0.775
defaults:
- training_hyperparams: coco2017_yolo_nas_pose_train_params
- dataset_params: coco_pose_estimation_yolo_nas_mosaic_dataset_params
- arch_params: yolo_nas_pose_s_arch_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
architecture: yolo_nas_pose_s
multi_gpu: DDP
num_gpus: 8
experiment_suffix: ""
experiment_name: coco2017_${architecture}_${experiment_suffix}_${dataset_params.dataset_params_suffix}
arch_params:
num_classes: ${dataset_params.num_joints}
checkpoint_params:
# For training Yolo-NAS-S pose estimation model we use pretrained weights for Yolo-NAS-S object detection model.
# By setting strict_load: key_matching we load only those weights that match the keys of the model.
checkpoint_path: https://sghub.deci.ai/models/yolo_nas_s_coco.pth
strict_load:
_target_: super_gradients.training.sg_trainer.StrictLoad
value: key_matching
dataset_params:
mosaic_prob: 0.5
train_dataloader_params:
batch_size: 48
val_dataloader_params:
batch_size: 48
---
Src/Super Gradients/Recipes/Coco2017 Yolo Nas S.Yaml (src/super_gradients/recipes/coco2017_yolo_nas_s.yaml)
YoloNAS-S Detection training on COCO2017 Dataset:
This training recipe is for demonstration purposes only. Pretrained models were trained using a different recipe.
So it will not be possible to reproduce the results of the pretrained models using this recipe.
Instructions:
0. Make sure that the data is stored in dataset_params.dataset_dir or add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below (feel free to check ReadMe)
1. Move to the project root (where you will find the ReadMe and src folder)
2. Run the command you want:
yolo_nas_s: python src/super_gradients/examples/train_from_recipe_example/train_from_recipe.py --config-name=coco2017_yolo_nas_s
#
defaults:
- training_hyperparams: coco2017_yolo_nas_train_params
- dataset_params: coco_detection_yolo_nas_dataset_params
- arch_params: yolo_nas_s_arch_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
train_dataloader: coco2017_train_yolo_nas
val_dataloader: coco2017_val_yolo_nas
load_checkpoint: False
resume: False
dataset_params:
train_dataloader_params:
batch_size: 32
arch_params:
num_classes: 80
training_hyperparams:
resume: ${resume}
mixed_precision: True
architecture: yolo_nas_s
multi_gpu: DDP
num_gpus: 8
experiment_suffix: ""
experiment_name: coco2017_${architecture}${experiment_suffix}
---
Src/Super Gradients/Recipes/Coco2017 Yolox.Yaml (src/super_gradients/recipes/coco2017_yolox.yaml)
YoloX Detection training on CoCo2017 Dataset:
YoloX trained in 640x640
Checkpoints + tensorboards: https://deci-pretrained-models.s3.amazonaws.com/yolox_coco/
Recipe runs with batch size = 16 X 8 gpus = 128.
Instructions:
0. Make sure that the data is stored in dataset_params.dataset_dir or add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below (feel free to check ReadMe)
1. Move to the project root (where you will find the ReadMe and src folder)
2. Run the command you want:
yolox_n: python -m super_gradients.train_from_recipe --config-name=coco2017_yolox architecture=yolox_n
yolox_t: python -m super_gradients.train_from_recipe --config-name=coco2017_yolox architecture=yolox_t
yolox_s: python -m super_gradients.train_from_recipe --config-name=coco2017_yolox architecture=yolox_s
yolox_m: python -m super_gradients.train_from_recipe --config-name=coco2017_yolox architecture=yolox_m
yolox_l: python -m super_gradients.train_from_recipe --config-name=coco2017_yolox architecture=yolox_l
yolox_x: python -m super_gradients.train_from_recipe --config-name=coco2017_yolox architecture=yolox_x
#
Training times and accuracies ([email protected] (COCO API, confidence 0.001, IoU threshold 0.6, test on 640x640 images):
yolox_n: 1d 16h 33m 9s on 8 NVIDIA GeForce RTX 3090, mAP: 26.77
yolox_t: 20h 43m 37s on 8 NVIDIA RTX A5000, mAP: 37.18
yolox_s: 1d 17h 40m 30s on 8 NVIDIA RTX A5000, mAP: 40.47
yolox_m: 1d 22h 23m 43s on 8 NVIDIA GeForce RTX 3090, mAP: 46.40
yolox_l: 2d 14h 11m 41s on 8 NVIDIA GeForce RTX 3090, mAP: 49.25
#
Using FAST LOSS
Training times and accuracies ([email protected] (COCO API, confidence 0.001, IoU threshold 0.6, test on 640x640 images):
yolox_n: COMING SOON
yolox_t: COMING SOON
yolox_s: 18h 23m 4s on 8 NVIDIA RTX A5000, mAP: 40.55
yolox_m: COMING SOON
yolox_l: COMING SOON
defaults:
- training_hyperparams: coco2017_yolox_train_params
- dataset_params: coco_detection_dataset_params
- arch_params: yolox_s_arch_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
train_dataloader: coco2017_train
val_dataloader: coco2017_val
load_checkpoint: False
resume: False
training_hyperparams:
resume: ${resume}
architecture: yolox_s
multi_gpu: DDP
num_gpus: 8
experiment_suffix: res${dataset_params.train_dataset_params.input_dim}
experiment_name: ${architecture}_coco2017_${experiment_suffix}
---
Src/Super Gradients/Recipes/Coco Segmentation Shelfnet Lw.Yaml (src/super_gradients/recipes/coco_segmentation_shelfnet_lw.yaml)
Shelfnet34_lw recipe for COCO segmentation 21 classes from PASCAL.
Reaches ~65.1 mIOU
Trained using 4 X 2080 Ti using DDP- takes ~ 2d 7h with batch size of 8 and batch accumulate of 3 (i.e effective batch
size is 483 = 96)
Logs and tensorboards: s3://deci-pretrained-models/shelfnet34_coco_segmentation_tensorboard/
Instructions:
0. Make sure that the data is stored in dataset_params.dataset_dir or add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below (feel free to check ReadMe)
1. Move to the project root (where you will find the ReadMe and src folder)
2. Run the command:
python -m super_gradients.train_from_recipe --config-name=coco_segmentation_shelfnet_lw
/!\ THIS RECIPE IS NOT SUPPORTED AT THE MOMENT /!\
defaults:
- training_hyperparams: coco_segmentation_shelfnet_lw_train_params
- dataset_params: coco_segmentation_dataset_params
- arch_params: shelfnet34_lw_arch_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
train_dataloader: coco_segmentation_train
val_dataloader: coco_segmentation_val
checkpoint_params:
strict_load: True
load_backbone: False
checkpoint_path:
resume: False
training_hyperparams:
resume: ${resume}
experiment_name: coco_segmentation_21_subclass_shelfnet34
multi_gpu: DDP
num_gpus: 4
architecture: shelfnet34_lw
---
Src/Super Gradients/Recipes/Imagenet Efficientnet.Yaml (src/super_gradients/recipes/imagenet_efficientnet.yaml)
Efficientnet-B0 Imagenet training
This example trains with effective batch size = 64 * 4 gpus = 256.
Epoch time on 4 X 3090Ti distributed training is ~ 16:25 minutes
Logs and tensorboards: s3://deci-pretrained-models/efficientnet_b0/
Instructions:
0. Make sure that the data is stored in dataset_params.dataset_dir or add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below (feel free to check ReadMe)
1. Move to the project root (where you will find the ReadMe and src folder)
2. Run the command:
python -m super_gradients.train_from_recipe --config-name=imagenet_efficientnet
defaults:
- training_hyperparams: imagenet_efficientnet_train_params
- dataset_params: imagenet_efficientnet_dataset_params
- arch_params: efficientnet_b0_arch_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
arch_params:
num_classes: 1000
train_dataloader: imagenet_train
val_dataloader: imagenet_val
resume: False
training_hyperparams:
resume: ${resume}
experiment_name: efficientnet_b0_imagenet
multi_gpu: DDP
num_gpus: 4
architecture: efficientnet_b0
---
Src/Super Gradients/Recipes/Imagenet Mobilenetv2.Yaml (src/super_gradients/recipes/imagenet_mobilenetv2.yaml)
MobilNetV2 ImageNetDataset training recipe.
Top1-Accuracy: 73.08
Learning rate and batch size parameters, using 2 GPUs with DDP:
initial_lr: 0.032 batch-size: 256 * 2gpus = 512
#
Instructions:
0. Make sure that the data is stored in dataset_params.dataset_dir or add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below (feel free to check ReadMe)
1. Move to the project root (where you will find the ReadMe and src folder)
2. Run the command:
python -m super_gradients.train_from_recipe --config-name=imagenet_mobilenetv2
defaults:
- training_hyperparams: imagenet_mobilenetv2_train_params
- dataset_params: imagenet_mobilenetv2_dataset_params
- arch_params: mobilenet_v2_arch_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
train_dataloader: imagenet_train
val_dataloader: imagenet_val
arch_params:
num_classes: 1000
dropout: 0.2
data_loader_num_workers: 8
resume: False
training_hyperparams:
resume: ${resume}
experiment_name: mobileNetv2_training
multi_gpu: DDP
num_gpus: 2
architecture: mobilenet_v2
---
Src/Super Gradients/Recipes/Imagenet Mobilenetv3 Base.Yaml (src/super_gradients/recipes/imagenet_mobilenetv3_base.yaml)
TODO: PRODUCE RESULTS AND ADD TENSORBOARDS, LOGS, TRAINING TIME ETC.
defaults:
- training_hyperparams: imagenet_mobilenetv3_train_params
- dataset_params: imagenet_mobilenetv3_dataset_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
train_dataloader: imagenet_train
val_dataloader: imagenet_val
resume: False
training_hyperparams:
resume: ${resume}
experiment_name: mobileNetv3_large_training
multi_gpu: DDP
num_gpus: 2
architecture: mobilenet_v3_large
---
Src/Super Gradients/Recipes/Imagenet Mobilenetv3 Large.Yaml (src/super_gradients/recipes/imagenet_mobilenetv3_large.yaml)
MobileNetV3 Large Imagenet classification training:
TODO: Add metrics
#
Instructions:
0. Make sure that the data is stored in dataset_params.dataset_dir or add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below (feel free to check ReadMe)
1. Move to the project root (where you will find the ReadMe and src folder)
2. Run the command:
python -m super_gradients.train_from_recipe --config-name=imagenet_mobilenetv3_large
defaults:
- imagenet_mobilenetv3_base
- arch_params: mobilenet_v3_large_arch_params
- _self_
- variable_setup
arch_params:
num_classes: 1000
dropout: 0.2
experiment_name: mobileNetv3_large_training
architecture: mobilenet_v3_large
---
Src/Super Gradients/Recipes/Imagenet Mobilenetv3 Small.Yaml (src/super_gradients/recipes/imagenet_mobilenetv3_small.yaml)
MobileNetV3 Small Imagenet classification training:
TODO: Add metrics
#
Instructions:
0. Make sure that the data is stored in dataset_params.dataset_dir or add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below (feel free to check ReadMe)
1. Move to the project root (where you will find the ReadMe and src folder)
2. Run the command:
python -m super_gradients.train_from_recipe --config-name=imagenet_mobilenetv3_small
defaults:
- imagenet_mobilenetv3_base
- arch_params: mobilenet_v3_small_arch_params
- _self_
- variable_setup
arch_params:
num_classes: 1000
dropout: 0.2
experiment_name: mobileNetv3_small_training
architecture: mobilenet_v3_small
---
Src/Super Gradients/Recipes/Imagenet RegnetY.Yaml (src/super_gradients/recipes/imagenet_regnetY.yaml)
RegnetY Imagenet classification training:
This example trains with batch_size = 256
Training time on a single GeForce RTX 2080 Ti, and top1 accuracies:
11 days for RegnetY200, 70.88
12 days for RegnetY400, 74.74
19 days for RegnetY600, 76.18
20 days for RegnetY800, 77.07
NOTE: Training should probably be lower as resources were shared among the above runs.
#
Logs and tensorboards at:
https://deci-pretrained-models.s3.amazonaws.com/RegnetY800/
https://deci-pretrained-models.s3.amazonaws.com/RegnetY600/
https://deci-pretrained-models.s3.amazonaws.com/RegnetY400/
https://deci-pretrained-models.s3.amazonaws.com/RegnetY200/
#
Instructions:
0. Make sure that the data is stored in dataset_params.dataset_dir or add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below (feel free to check ReadMe)
1. Move to the project root (where you will find the ReadMe and src folder)
2. Run the command:
regnetY200: python -m super_gradients.train_from_recipe --config-name=imagenet_regnetY architecture=regnetY200
regnetY400: python -m super_gradients.train_from_recipe --config-name=imagenet_regnetY architecture=regnetY400
regnetY600: python -m super_gradients.train_from_recipe --config-name=imagenet_regnetY architecture=regnetY600
regnetY800: python -m super_gradients.train_from_recipe --config-name=imagenet_regnetY architecture=regnetY800
defaults:
- training_hyperparams: imagenet_regnetY_train_params
- dataset_params: imagenet_regnetY_dataset_params
- arch_params: regnetY_arch_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
arch_params:
num_classes: 1000
dropout_prob: 0.5
droppath_prob: 0.0
train_dataloader: imagenet_train
val_dataloader: imagenet_val
load_checkpoint: False
resume: False
training_hyperparams:
resume: ${resume}
multi_gpu: Off
num_gpus: 1
architecture: regnetY800
experiment_name: ${architecture}
---
Src/Super Gradients/Recipes/Imagenet Repvgg.Yaml (src/super_gradients/recipes/imagenet_repvgg.yaml)
RepVGGA0 Imagenet classification training:
This example trains with batch_size = 64 * 4 GPUs, total 256.
Training time on 4 X GeForce RTX 3090 Ti is 10min / epoch, total time ~ 20h 22m (DistributedDataParallel).
Reach => 72.05 Top1 accuracy.
#
Log and tensorboard at s3://deci-pretrained-models/repvggg-a0-imagenet-tensorboard/
#
Instructions:
0. Make sure that the data is stored in dataset_params.dataset_dir or add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below (feel free to check ReadMe)
1. Move to the project root (where you will find the ReadMe and src folder)
2. Run the command:
python -m super_gradients.train_from_recipe --config-name=imagenet_repvgg
defaults:
- training_hyperparams: imagenet_repvgg_train_params
- dataset_params: imagenet_dataset_params
- arch_params: repvgg_arch_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
arch_params:
num_classes: 1000
build_residual_branches: True
train_dataloader: imagenet_train
val_dataloader: imagenet_val
resume: False
training_hyperparams:
resume: ${resume}
experiment_name: repvgg_a0_imagenet_reproduce_fix
multi_gpu: DDP
num_gpus: 4
architecture: repvgg_a0
---
Src/Super Gradients/Recipes/Imagenet Resnet50.Yaml (src/super_gradients/recipes/imagenet_resnet50.yaml)
ResNet50 Imagenet classification training:
This example trains with batch_size = 64 * 4 GPUs, total 256.
Training time on 4 x GeForce RTX A5000 is 15min / epoch.
Reach => 79.47 Top1 accuracy.
#
Log and tensorboard at s3://deci-pretrained-models/ResNet50_ImageNet/average_model.pth
Instructions:
0. Make sure that the data is stored in dataset_params.dataset_dir or add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below (feel free to check ReadMe)
1. Move to the project root (where you will find the ReadMe and src folder)
2. Run the command:
python -m super_gradients.train_from_recipe --config-name=imagenet_resnet50
defaults:
- training_hyperparams: imagenet_resnet50_train_params
- dataset_params: imagenet_resnet50_dataset_params
- arch_params: resnet50_arch_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
arch_params:
droppath_prob: 0.05
train_dataloader: imagenet_train
val_dataloader: imagenet_val
resume: False
training_hyperparams:
resume: ${resume}
experiment_name: resnet50_imagenet
multi_gpu: DDP
num_gpus: 4
architecture: resnet50
---
Src/Super Gradients/Recipes/Imagenet Resnet50 Kd.Yaml (src/super_gradients/recipes/imagenet_resnet50_kd.yaml)
ResNet50 Imagenet classification training:
This example trains with batch_size = 192 * 8 GPUs, total 1536.
Training time on 8 x GeForce RTX A5000 is 9min / epoch.
Reach => 81.91 Top1 accuracy.
#
Log and tensorboard at s3://deci-pretrained-models/KD_ResNet50_Beit_Base_ImageNet/average_model.pth
Instructions:
0. Make sure that the data is stored in dataset_params.dataset_dir or add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below (feel free to check ReadMe)
1. Move to the project root (where you will find the ReadMe and src folder)
2. Run the command:
python -m super_gradients.train_from_kd_recipe --config-name=imagenet_resnet50_kd
defaults:
- training_hyperparams: imagenet_resnet50_kd_train_params
- dataset_params: imagenet_resnet50_kd_dataset_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
train_dataloader: imagenet_train
val_dataloader: imagenet_val
resume: False
training_hyperparams:
resume: ${resume}
loss: KDLogitsLoss
criterion_params:
distillation_loss_coeff: 0.8
task_loss_fn:
_target_: super_gradients.training.losses.label_smoothing_cross_entropy_loss.CrossEntropyLoss
arch_params:
teacher_input_adapter:
_target_: super_gradients.training.utils.kd_trainer_utils.NormalizationAdapter
mean_original: [0.485, 0.456, 0.406]
std_original: [0.229, 0.224, 0.225]
mean_required: [0.5, 0.5, 0.5]
std_required: [0.5, 0.5, 0.5]
student_arch_params:
num_classes: 1000
teacher_arch_params:
num_classes: 1000
image_size: [224, 224]
patch_size: [16, 16]
teacher_checkpoint_params:
load_backbone: False # whether to load only backbone part of checkpoint
checkpoint_path: # checkpoint path that is not located in super_gradients/checkpoints
strict_load: # key matching strictness for loading checkpoint's weights
_target_: super_gradients.training.sg_trainer.StrictLoad
value: True
pretrained_weights: imagenet
checkpoint_params:
teacher_pretrained_weights: imagenet
student_checkpoint_params:
load_backbone: False # whether to load only backbone part of checkpoint
checkpoint_path: # checkpoint path that is not located in super_gradients/checkpoints
strict_load: # key matching strictness for loading checkpoint's weights
_target_: super_gradients.training.sg_trainer.StrictLoad
value: True
pretrained_weights: # a string describing the dataset of the pretrained weights (for example "imagenent").
run_teacher_on_eval: True
experiment_name: resnet50_imagenet_KD_Model
multi_gpu: DDP
num_gpus: 8
architecture: kd_module
student_architecture: resnet50
teacher_architecture: beit_base_patch16_224
---
Src/Super Gradients/Recipes/Imagenet Vit Base.Yaml (src/super_gradients/recipes/imagenet_vit_base.yaml)
ViT Imagenet1K fine tuning from Imagenet21K classification training:
This example trains with batch_size = 64 * 8 GPUs, total 512.
Training time on 8 x GeForce RTX A5000 is 15min / epoch.
ViT base : 84.15
#
Log and tensorboard at s3://deci-pretrained-models/vit_base_imagenet1k/
Instructions:
0. Make sure that the data is stored in dataset_params.dataset_dir or add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below (feel free to check ReadMe)
1. Move to the project root (where you will find the ReadMe and src folder)
2. Run the command:
python -m super_gradients.train_from_recipe --config-name=imagenet_vit_base
defaults:
- training_hyperparams: imagenet_vit_train_params
- dataset_params: imagenet_vit_base_dataset_params
- arch_params: vit_base_arch_params
- checkpoint_params: vit_base_imagenet_checkpoint_params
- _self_
- variable_setup
train_dataloader: imagenet_train
val_dataloader: imagenet_val
resume: False
training_hyperparams:
resume: ${resume}
experiment_name: vit_base_imagenet1k
architecture: vit_base
multi_gpu: DDP
num_gpus: 8
---
Src/Super Gradients/Recipes/Imagenet Vit Large.Yaml (src/super_gradients/recipes/imagenet_vit_large.yaml)
ViT Imagenet1K fine tuning from Imagenet21K classification training:
This example trains with batch_size = 32 * 8 GPUs, total 256.
Training time on 8 x GeForce RTX A5000 is 52min / epoch.
ViT Large : 85.64 (Final averaged model)
#
Log and tensorboard at s3://deci-pretrained-models/vit_large_cutmix_randaug_v2_lr=0.03/
Instructions:
0. Make sure that the data is stored in dataset_params.dataset_dir or add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below (feel free to check ReadMe)
1. Move to the project root (where you will find the ReadMe and src folder)
2. Run the command:
python -m super_gradients.train_from_recipe --config-name=imagenet_vit_large
defaults:
- imagenet_vit_base
- _self_
- variable_setup
dataset_params:
train_dataloader_params:
batch_size: 32
training_hyperparams:
initial_lr: 0.06
average_best_models: True
architecture: vit_large
experiment_name: vit_large_imagenet1k
multi_gpu: DDP
num_gpus: 8
---
Src/Super Gradients/Recipes/Roboflow Ppyoloe.Yaml (src/super_gradients/recipes/roboflow_ppyoloe.yaml)
Checkout the datasets at https://universe.roboflow.com/roboflow-100?ref=blog.roboflow.com
#
dataset_name refers to the official name of the dataset.
You can find it in the url of the dataset: https://universe.roboflow.com/roboflow-100/digits-t2eg6 -> digits-t2eg6
#
Example: python -m super_gradients.train_from_recipe --config-name=roboflow_ppyoloe dataset_name=digits-t2eg6
defaults:
- training_hyperparams: coco2017_ppyoloe_train_params
- dataset_params: roboflow_detection_dataset_params
- checkpoint_params: default_checkpoint_params
- arch_params: ppyoloe_m_arch_params
- _self_
- variable_setup
dataset_name: ??? # Placeholder for the name of the dataset you want to use (e.g. "digits-t2eg6")
dataset_params:
dataset_name: ${dataset_name}
num_classes: ${roboflow_dataset_num_classes:${dataset_name}}
architecture: ppyoloe_m
arch_params:
num_classes: ${num_classes}
train_dataloader: roboflow_train_yolox
val_dataloader: roboflow_val_yolox
load_checkpoint: False
checkpoint_params:
pretrained_weights: coco
result_path: # By defaults saves results in checkpoints directory
resume: False
training_hyperparams:
resume: ${resume}
max_epochs: 100
mixed_precision: True
criterion_params:
num_classes: ${num_classes}
reg_max: ${arch_params.head.reg_max}
phase_callbacks:
- RoboflowResultCallback:
dataset_name: ${dataset_name}
output_path: ${result_path}
loss: PPYoloELoss
valid_metrics_list:
- DetectionMetrics:
score_thres: 0.1
top_k_predictions: 300
num_cls: ${num_classes}
normalize_targets: True
post_prediction_callback:
_target_: super_gradients.training.models.detection_models.pp_yolo_e.PPYoloEPostPredictionCallback
score_threshold: 0.01
nms_top_k: 1000
max_predictions: 300
nms_threshold: 0.7
multi_gpu: DDP
num_gpus:
experiment_name: ${architecture}_roboflow_${dataset_name}
---
Src/Super Gradients/Recipes/Roboflow Yolo Nas M.Yaml (src/super_gradients/recipes/roboflow_yolo_nas_m.yaml)
A recipe to fine-tune YoloNAS on Roboflow datasets.
Checkout the datasets at https://universe.roboflow.com/roboflow-100?ref=blog.roboflow.com
#
dataset_name refers to the official name of the dataset.
You can find it in the url of the dataset: https://universe.roboflow.com/roboflow-100/digits-t2eg6 -> digits-t2eg6
#
Example: python -m super_gradients.train_from_recipe --config-name=roboflow_yolo_nas_m dataset_name=digits-t2eg6
defaults:
- training_hyperparams: coco2017_yolo_nas_train_params
- dataset_params: roboflow_detection_dataset_params
- checkpoint_params: default_checkpoint_params
- arch_params: yolo_nas_m_arch_params
- _self_
- variable_setup
train_dataloader: roboflow_train_yolox
val_dataloader: roboflow_val_yolox
dataset_name: ??? # Placeholder for the name of the dataset you want to use (e.g. "digits-t2eg6")
dataset_params:
dataset_name: ${dataset_name}
train_dataloader_params:
batch_size: 12
val_dataloader_params:
batch_size: 16
num_classes: ${roboflow_dataset_num_classes:${dataset_name}}
architecture: yolo_nas_m
arch_params:
num_classes: ${num_classes}
load_checkpoint: False
checkpoint_params:
pretrained_weights: coco
result_path: # By defaults saves results in checkpoints directory
resume: False
training_hyperparams:
resume: ${resume}
zero_weight_decay_on_bias_and_bn: True
lr_warmup_epochs: 3
warmup_mode: LinearEpochLRWarmup
initial_lr: 4e-4
cosine_final_lr_ratio: 0.1
optimizer_params:
weight_decay: 0.0001
ema: True
ema_params:
decay: 0.9
max_epochs: 100
mixed_precision: True
criterion_params:
num_classes: ${num_classes}
reg_max: 16
phase_callbacks: []
loss: PPYoloELoss
valid_metrics_list:
- DetectionMetrics_050:
score_thres: 0.1
top_k_predictions: 300
num_cls: ${num_classes}
normalize_targets: True
post_prediction_callback:
_target_: super_gradients.training.models.detection_models.pp_yolo_e.PPYoloEPostPredictionCallback
score_threshold: 0.01
nms_top_k: 1000
max_predictions: 300
nms_threshold: 0.7
metric_to_watch: '[email protected]'
multi_gpu: Off
num_gpus: 1
experiment_suffix: ""
experiment_name: ${architecture}_roboflow_${dataset_name}${experiment_suffix}
---
Src/Super Gradients/Recipes/Roboflow Yolo Nas S.Yaml (src/super_gradients/recipes/roboflow_yolo_nas_s.yaml)
A recipe to fine-tune YoloNAS on Roboflow datasets.
Checkout the datasets at https://universe.roboflow.com/roboflow-100?ref=blog.roboflow.com
#
dataset_name refers to the official name of the dataset.
You can find it in the url of the dataset: https://universe.roboflow.com/roboflow-100/digits-t2eg6 -> digits-t2eg6
#
Example: python -m super_gradients.train_from_recipe --config-name=roboflow_yolo_nas_s dataset_name=digits-t2eg6
defaults:
- training_hyperparams: coco2017_yolo_nas_train_params
- dataset_params: roboflow_detection_dataset_params
- checkpoint_params: default_checkpoint_params
- arch_params: yolo_nas_s_arch_params
- _self_
- variable_setup
train_dataloader: roboflow_train_yolox
val_dataloader: roboflow_val_yolox
dataset_name: ??? # Placeholder for the name of the dataset you want to use (e.g. "digits-t2eg6")
dataset_params:
dataset_name: ${dataset_name}
train_dataloader_params:
batch_size: 16
val_dataloader_params:
batch_size: 16
num_classes: ${roboflow_dataset_num_classes:${dataset_name}}
architecture: yolo_nas_s
arch_params:
num_classes: ${num_classes}
load_checkpoint: False
checkpoint_params:
pretrained_weights: coco
result_path: # By defaults saves results in checkpoints directory
resume: False
training_hyperparams:
resume: ${resume}
zero_weight_decay_on_bias_and_bn: True
lr_warmup_epochs: 3
warmup_mode: LinearEpochLRWarmup
initial_lr: 5e-4
cosine_final_lr_ratio: 0.1
optimizer_params:
weight_decay: 0.0001
ema: True
ema_params:
decay: 0.9
max_epochs: 100
mixed_precision: True
criterion_params:
num_classes: ${num_classes}
reg_max: 16
phase_callbacks: []
loss: PPYoloELoss
valid_metrics_list:
- DetectionMetrics_050:
score_thres: 0.1
top_k_predictions: 300
num_cls: ${num_classes}
normalize_targets: True
post_prediction_callback:
_target_: super_gradients.training.models.detection_models.pp_yolo_e.PPYoloEPostPredictionCallback
score_threshold: 0.01
nms_top_k: 1000
max_predictions: 300
nms_threshold: 0.7
metric_to_watch: '[email protected]'
multi_gpu: Off
num_gpus: 1
experiment_suffix: ""
experiment_name: ${architecture}_roboflow_${dataset_name}${experiment_suffix}
---
Src/Super Gradients/Recipes/Roboflow Yolo Nas S Qat.Yaml (src/super_gradients/recipes/roboflow_yolo_nas_s_qat.yaml)
defaults:
- roboflow_yolo_nas_s
- quantization_params: default_quantization_params
- _self_
checkpoint_params:
checkpoint_path: ???
strict_load: no_key_matching
pre_launch_callbacks_list:
- QATRecipeModificationCallback:
batch_size_divisor: 2
max_epochs_divisor: 10
lr_decay_factor: 0.01
warmup_epochs_divisor: 10
cosine_final_lr_ratio: 0.01
disable_phase_callbacks: True
disable_augmentations: False
---
Src/Super Gradients/Recipes/Roboflow Yolox.Yaml (src/super_gradients/recipes/roboflow_yolox.yaml)
Checkout the datasets at https://universe.roboflow.com/roboflow-100?ref=blog.roboflow.com
#
dataset_name refers to the official name of the dataset.
You can find it in the url of the dataset: https://universe.roboflow.com/roboflow-100/digits-t2eg6 -> digits-t2eg6
#
Example: python -m super_gradients.train_from_recipe --config-name=roboflow_yolox dataset_name=digits-t2eg6
defaults:
- training_hyperparams: coco2017_yolox_train_params
- dataset_params: roboflow_detection_dataset_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
dataset_name: ??? # Placeholder for the name of the dataset you want to use (e.g. "digits-t2eg6")
dataset_params:
dataset_name: ${dataset_name}
num_classes: ${roboflow_dataset_num_classes:${dataset_name}}
architecture: yolox_m
arch_params:
num_classes: ${num_classes}
yolo_type: 'yoloX'
depth_mult_factor: 0.67
width_mult_factor: 0.75
train_dataloader: roboflow_train_yolox
val_dataloader: roboflow_val_yolox
load_checkpoint: False
checkpoint_params:
pretrained_weights: coco
result_path: # By defaults saves results in checkpoints directory
resume: False
training_hyperparams:
max_epochs: 100
resume: ${resume}
criterion_params:
num_classes: ${num_classes}
train_metrics_list:
- DetectionMetrics:
normalize_targets: True
post_prediction_callback:
_target_: super_gradients.training.models.detection_models.yolo_base.YoloXPostPredictionCallback
iou: 0.65
conf: 0.01
num_cls: 80
valid_metrics_list:
- DetectionMetrics:
normalize_targets: True
post_prediction_callback:
_target_: super_gradients.training.models.detection_models.yolo_base.YoloXPostPredictionCallback
iou: 0.65
conf: 0.01
num_cls: 80
multi_gpu: DDP
num_gpus: 3
experiment_name: ${architecture}_roboflow_${dataset_name}
---
Src/Super Gradients/Recipes/Script Generate Rescoring Data Dekr Coco2017.Yaml (src/super_gradients/recipes/script_generate_rescoring_data_dekr_coco2017.yaml)
This script contains a recipe to generate the training data for rescoring net for DEKR architecutre on COCO2017 dataset.
See documentation/source/PoseEstimation.md for more details.
#
Example usage:
python -m super_gradients.scripts.generate_rescoring_training_data --config-name=script_generate_rescoring_data_dekr_coco2017 rescoring_data_dir=WHERE_TO_STORE_FILES checkpoint=PATH_TO_TRAINED_MODEL_CHECKPOINT.
defaults:
- dataset_params: coco_pose_estimation_dekr_dataset_params
- checkpoint_params: default_checkpoint_params
- arch_params: pose_dekr_w32_no_dc_arch_params
- _self_
architecture: dekr_w32_no_dc
rescoring_data_dir: ???
train_dataloader: coco2017_pose_train
val_dataloader: coco2017_pose_val
arch_params:
num_classes: ${dataset_params.num_joints}
checkpoint_params:
checkpoint_path: # Put the path to the checkpoint here
post_prediction_callback:
_target_: super_gradients.training.utils.pose_estimation.DEKRPoseEstimationDecodeCallback
max_num_people: 30
keypoint_threshold: 0.05
nms_threshold: 0.05
nms_num_threshold: 8
output_stride: 4
apply_sigmoid: False # We use Flip-TTA and apply sigmoid there
THE FOLLOWING PARAMS ARE DIRECTLY USED BY HYDRA
hydra:
run:
# Set the output directory (i.e. where .hydra folder that logs all the input params will be generated)
dir: .
---
Src/Super Gradients/Recipes/Supervisely Unet.Yaml (src/super_gradients/recipes/supervisely_unet.yaml)
Binary segmentation training example of UNet model on the Supervisely person dataset.
Instructions:
0. Make sure that the data is stored in dataset_params.dataset_dir or add "dataset_params.data_dir=<PATH-TO-DATASET>" at the end of the command below (feel free to check ReadMe)
1. Move to the project root (where you will find the ReadMe and src folder)
2. Run the command:
UNet: python -m super_gradients.train_from_recipe --config-name=supervisely_unet
#
Validation Target (Person class) IoU and training time:
UNet: input-size: [480, 320] mIoU: 89.18 1 X RTX A5000, 4 H
#
Logs, tensorboards and network checkpoints:
UNet: https://deci-pretrained-models.s3.amazonaws.com/unet/supervisely/
#
defaults:
- training_hyperparams: supervisely_default_train_params
- dataset_params: supervisely_persons_dataset_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
architecture: unet
arch_params:
num_classes: 1
use_aux_heads: False
training_hyperparams:
initial_lr: 0.025
loss:
BCEDiceLoss:
loss_weights: [ 1., 1. ]
logits: True
dataset_params:
batch_size: 16
multi_gpu: OFF
experiment_name: unet_supervisely
---
Src/Super Gradients/Recipes/User Recipe Mnist As External Dataset Example.Yaml (src/super_gradients/recipes/user_recipe_mnist_as_external_dataset_example.yaml)
The purpose of the example below is to demonstrate the use of registry for external objects for training.
- We train mobilenet_v2 on a user dataset which is not defined in ALL_DATASETS using the dataloader registry.
- We leverage predefined configs from cifar_10 training recipe in our repo.
#
In order for the registry to work, we must trigger the registry of the user's objects by importing their module at
the top of the training script. Hence, we created a similar script to our classic train_from_recipe but with the imports
on top. Once imported, all the registry decorated objects will be resolved (i.e user_mnist_train will be resolved
to the dataloader of our user's)
defaults:
- training_hyperparams: cifar10_resnet_train_params
- dataset_params: cifar10_dataset_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
arch_params:
num_classes: 10
in_channels: 1
dataset_params:
train_dataset_params:
root: ./data/mnist
train: True
transforms:
- RandomHorizontalFlip
- ToTensor
target_transform: null
download: True
train_dataloader_params:
dataset: MnistDataset
batch_size: 256
num_workers: 8
drop_last: False
pin_memory: True
val_dataset_params:
root: ./data/mnist
train: False
transforms:
- ToTensor
target_transform: null
download: True
val_dataloader_params:
dataset: MnistDataset
batch_size: 512
num_workers: 8
drop_last: False
pin_memory: True
resume: False
training_hyperparams:
resume: ${resume}
max_epochs: 3
architecture: mobilenet_v2
experiment_name: mobilenet_v2_mnist
---
Src/Super Gradients/Recipes/User Recipe Mnist Example.Yaml (src/super_gradients/recipes/user_recipe_mnist_example.yaml)
The purpose of the example below is to demonstrate the use of registry for external objects for training.
- We train mobilenet_v2 on a user dataset which is not defined in ALL_DATASETS using the dataloader registry.
- We leverage predefined configs from cifar_10 training recipe in our repo.
#
In order for the registry to work, we must trigger the registry of the user's objects by importing their module at
the top of the training script. Hence, we created a similar script to our classic train_from_recipe but with the imports
on top. Once imported, all the registry decorated objects will be resolved (i.e user_mnist_train will be resolved
to the dataloader of our user's)
#
Differently from user_recipe_mnist_example, here we demonstrate how to use train_from_recipe, without the need to implement a DataLoader class for registry.
Instead- we work straight with the user defined datasets, which is the simpler option when one does not need their own DataLoader implementation.
We do so by Dropping the train_datalaoder, valid_dataloader fields from the recipe's config, while specifying the dataset arg in
train_dataloader_params, valid_dataloader_params.
defaults:
- training_hyperparams: cifar10_resnet_train_params
- dataset_params: cifar10_dataset_params
- checkpoint_params: default_checkpoint_params
- _self_
- variable_setup
arch_params:
num_classes: 10
in_channels: 1
dataset_params:
train_dataset_params:
root: ./data/mnist
train: True
transforms:
- RandomHorizontalFlip
- ToTensor
target_transform: null
download: True
train_dataloader_params:
dataset: Mnist
batch_size: 256
num_workers: 8
drop_last: False
pin_memory: True
val_dataset_params:
root: ./data/mnist
train: False
transforms:
- ToTensor
target_transform: null
download: True
val_dataloader_params:
batch_size: 512
num_workers: 8
drop_last: False
pin_memory: True
resume: False
training_hyperparams:
resume: ${resume}
max_epochs: 3
architecture: mobilenet_v2
experiment_name: mobilenet_v2_mnist
---
Src/Super Gradients/Recipes/Variable Setup.Yaml (src/super_gradients/recipes/variable_setup.yaml)
Varaible setup for shortcuts and setting the hydra output directory.
Any SG Recipe should set this yaml file as a default, after _self_, i.e at the top of your recipe file:
#
defaults:
- training_hyperparams: my_train_params
- dataset_params: my_dataset_params
- arch_params: my_arch_params
- checkpoint_params: my_checkpoint_params
- _self_
- variable_setup
#
#
Interpolates the shortcuts defined below, with their aliases (see comments near each parameter).
When any of the above are not set, they will be populated with the original values (for example
config.lr will be set with config.training_hyperparams.initial_lr) for clarity in logs.
#
In other words, the following training launch commands are equivalent:
#
python train_from_recipe --config-name=recipe lr=0.003
#
python train_from_recipe --config-name=recipe config.training_hyperparams.initial_lr=0.003
#
Note that interpolation is done by triggering RecipeShortcutsCallbackm which is a Hydra Callback (see http://hydra.cc/docs/experimental/callbacks/)
so interpolation of these in other yaml configuration files won't be present.
lr: # config.training_hyperparams.initial_lr
batch_size: # config.dataset_params.train_dataloader_params.batch_size
val_batch_size: # config.dataset_params.val_dataloader_params.batch_size
ema: # config.training_hyperparams.ema
epochs: # config.training_hyperparams.max_epochs
resume: # config.training_hyperparams.resume
num_workers: # config.dataset_params.train_dataloader_params.num_workers and config.dataset_params.val_dataloader_params.num_workers
ckpt_root_dir:
THE FOLLOWING PARAMS ARE DIRECTLY USED BY HYDRA
hydra:
callbacks:
shortcuts_cb:
_target_: super_gradients.common.environment.omegaconf_utils.RecipeShortcutsCallback
run:
# Set the output directory (i.e. where .hydra folder that logs all the input params will be generated)
dir: ${hydra_output_dir:${ckpt_root_dir}, ${experiment_name}}
---
Src/Super Gradients/Recipes/Anchors/Ssd Anchors.Yaml (src/super_gradients/recipes/anchors/ssd_anchors.yaml)
stride_N_plus is for models where the first skip begins from the feature map with output stride N and higher
grids of [input_size / N x input_size / N] and smaller
NOTE: changing anchors to a different stride requires updating output_paths in model anch params
because feat_size are hardcoded and won't automatically change in a model
256x256:
stride_16_plus:
_target_: super_gradients.training.utils.ssd_utils.DefaultBoxes
fig_size: 256
feat_size: [ 32, 16, 8, 4, 2, 1 ]
scales: [ 18, 38, 84, 131, 177, 223, 269 ]
aspect_ratios: [ [ 2 ], [ 2, 3 ], [ 2, 3 ], [ 2, 3 ], [ 2 ], [ 2 ] ]
scale_xy: 0.1
scale_wh: 0.2
stride_8_plus: [[2, 3], [2, 3], [2, 3], [2, 3], [2, 3], [2, 3]]
300x300:
stride_8_plus:
_target_: super_gradients.training.utils.ssd_utils.DefaultBoxes
fig_size: 300
feat_size: [38, 19, 10, 5, 3, 2]
scales: [21, 45, 99, 153, 207, 261, 315]
aspect_ratios: [[2], [2, 3], [2, 3], [2, 3], [2], [2]]
scale_xy: 0.1
scale_wh: 0.2
stride_16_plus:
_target_: super_gradients.training.utils.ssd_utils.DefaultBoxes
fig_size: 300
feat_size: [19, 10, 5, 3, 2, 1]
scales: [21, 45, 99, 153, 207, 261, 315]
aspect_ratios: [[2, 3], [2, 3], [2, 3], [2, 3], [2, 3], [2, 3]]
scale_xy: 0.1
scale_wh: 0.2
320x320:
stride_8_plus:
_target_: super_gradients.training.utils.ssd_utils.DefaultBoxes
fig_size: 320
feat_size: [ 40, 20, 10, 5, 3, 2 ]
scales: [ 22, 48, 106, 163, 221, 278, 336 ]
aspect_ratios: [ [ 2 ], [ 2, 3 ], [ 2, 3 ], [ 2, 3 ], [ 2 ], [ 2 ] ]
scale_xy: 0.1
scale_wh: 0.2
stride_16_plus:
_target_: super_gradients.training.utils.ssd_utils.DefaultBoxes
fig_size: 320
feat_size: [ 20, 10, 5, 3, 2, 1 ]
scales: [ 22, 48, 106, 163, 221, 278, 336 ]
aspect_ratios: [ [ 2, 3 ], [ 2, 3 ], [ 2, 3 ], [ 2, 3 ], [ 2, 3 ], [ 2, 3 ] ]
scale_xy: 0.1
scale_wh: 0.2
stride_16_plus_big:
_target_: super_gradients.training.utils.ssd_utils.DefaultBoxes
fig_size: 320
feat_size: [ 20, 10, 5, 3, 2, 1 ]
scales: [ 32, 82, 133, 184, 235, 285, 336 ]
aspect_ratios: [ [ 2, 3 ], [ 2, 3 ], [ 2, 3 ], [ 2, 3 ], [ 2, 3 ], [ 2, 3 ] ]
scale_xy: 0.1
scale_wh: 0.2
---
Src/Super Gradients/Recipes/Arch Params/Efficientnet Arch Params.Yaml (src/super_gradients/recipes/arch_params/efficientnet_arch_params.yaml)
backbone_mode: False # cuts off classification head
batch_norm_momentum: 0.99 # value used for the running_mean and running_var computation
batch_norm_epsilon: 1e-3 # value added to the denominator for numerical stability
image_size: # net's input size
see doc in super_gradients/training/models/efficientnet.py round_filters
width_coefficient:
depth_divisor: 8
min_depth:
depth_coefficient:
dropout_rate: # dropout probability in final layer
num_classes: # number of outputs of the classification head
drop_connect_rate: 0.2 # connection dropout probability
---
Src/Super Gradients/Recipes/Arch Params/Efficientnet B0 Arch Params.Yaml (src/super_gradients/recipes/arch_params/efficientnet_b0_arch_params.yaml)
image_size: 224
width_coefficient: 1.0
min_depth:
depth_coefficient: 1.0
dropout_rate: 0.2
num_classes:
---
Src/Super Gradients/Recipes/Arch Params/Efficientnet B1 Arch Params.Yaml (src/super_gradients/recipes/arch_params/efficientnet_b1_arch_params.yaml)
image_size: 240
width_coefficient: 1.0
min_depth:
depth_coefficient: 1.1
dropout_rate: 0.2
num_classes:
---
Src/Super Gradients/Recipes/Arch Params/Efficientnet B2 Arch Params.Yaml (src/super_gradients/recipes/arch_params/efficientnet_b2_arch_params.yaml)
image_size: 260
width_coefficient: 1.1
min_depth:
depth_coefficient: 1.2
dropout_rate: 0.3
num_classes:
---
Src/Super Gradients/Recipes/Arch Params/Efficientnet B3 Arch Params.Yaml (src/super_gradients/recipes/arch_params/efficientnet_b3_arch_params.yaml)
image_size: 300
width_coefficient: 1.2
min_depth:
depth_coefficient: 1.4
dropout_rate: 0.3
num_classes:
---
Src/Super Gradients/Recipes/Arch Params/Efficientnet B4 Arch Params.Yaml (src/super_gradients/recipes/arch_params/efficientnet_b4_arch_params.yaml)
image_size: 380
width_coefficient: 1.4
min_depth:
depth_coefficient: 1.8
dropout_rate: 0.4
num_classes:
---
Src/Super Gradients/Recipes/Arch Params/Efficientnet B5 Arch Params.Yaml (src/super_gradients/recipes/arch_params/efficientnet_b5_arch_params.yaml)
image_size: 456
width_coefficient: 1.6
min_depth:
depth_coefficient: 2.2
dropout_rate: 0.4
num_classes:
---
Src/Super Gradients/Recipes/Arch Params/Efficientnet B6 Arch Params.Yaml (src/super_gradients/recipes/arch_params/efficientnet_b6_arch_params.yaml)
image_size: 528
width_coefficient: 1.8
min_depth:
depth_coefficient: 2.6
dropout_rate: 0.5
num_classes:
---
Src/Super Gradients/Recipes/Arch Params/Efficientnet B7 Arch Params.Yaml (src/super_gradients/recipes/arch_params/efficientnet_b7_arch_params.yaml)
image_size: 600
width_coefficient: 2.0
min_depth:
depth_coefficient: 3.1
dropout_rate: 0.5
num_classes:
---
Src/Super Gradients/Recipes/Arch Params/Efficientnet B8 Arch Params.Yaml (src/super_gradients/recipes/arch_params/efficientnet_b8_arch_params.yaml)
image_size: 672
width_coefficient: 2.2
min_depth:
depth_coefficient: 3.6
dropout_rate: 0.5
num_classes:
---
Src/Super Gradients/Recipes/Arch Params/Efficientnet L2 Arch Params.Yaml (src/super_gradients/recipes/arch_params/efficientnet_l2_arch_params.yaml)
image_size: 800
width_coefficient: 4.3
min_depth:
depth_coefficient: 5.3
dropout_rate: 0.5
num_classes:
---
Src/Super Gradients/Recipes/Arch Params/Mobilenet V2 Arch Params.Yaml (src/super_gradients/recipes/arch_params/mobilenet_v2_arch_params.yaml)
structure: # model structure
num_classes: # number of outputs of the classification head
width_mult: # model's width multiplier
dropout: # dropout probability of classifier fully-connected layer
---
Src/Super Gradients/Recipes/Arch Params/Mobilenet V3 Arch Params.Yaml (src/super_gradients/recipes/arch_params/mobilenet_v3_arch_params.yaml)
structure: # model structure
mode: # mode (i.e backbone)
num_classes: # number of outputs of the classification head
width_mult: # model's width multiplier
---
Src/Super Gradients/Recipes/Arch Params/Mobilenet V3 Large Arch Params.Yaml (src/super_gradients/recipes/arch_params/mobilenet_v3_large_arch_params.yaml)
defaults:
- mobilenet_v3_arch_params
structure: [[3, 1, 16, 0, 0, 1],[3, 4, 24, 0, 0, 2],[3, 3, 24, 0, 0, 1],[5, 3, 40, 1, 0, 2],[5, 3, 40, 1, 0, 1], [5, 3, 40, 1, 0, 1],[3, 6, 80, 0, 1, 2], [3, 2.5, 80, 0, 1, 1], [3, 2.3, 80, 0, 1, 1], [3, 2.3, 80, 0, 1, 1],[3, 6, 112, 1, 1, 1],[3, 6, 112, 1, 1, 1],[5, 6, 160, 1, 1, 2], [5, 6, 160, 1, 1, 1], [5, 6, 160, 1, 1, 1]]
mode: large
num_classes:
width_mult: 1
---
Src/Super Gradients/Recipes/Arch Params/Mobilenet V3 Small Arch Params.Yaml (src/super_gradients/recipes/arch_params/mobilenet_v3_small_arch_params.yaml)
defaults:
- mobilenet_v3_arch_params
structure: [[3, 1, 16, 1, 0, 2], [ 3, 4.5, 24, 0, 0, 2 ], [ 3, 3.67, 24, 0, 0, 1 ], [ 5, 4, 40, 1, 1, 2 ], [ 5, 6, 40, 1, 1, 1 ], [ 5, 6, 40, 1, 1, 1 ], [ 5, 3, 48, 1, 1, 1 ], [ 5, 3, 48, 1, 1, 1 ], [ 5, 6, 96, 1, 1, 2 ], [ 5, 6, 96, 1, 1, 1 ], [ 5, 6, 96, 1, 1, 1 ]]
mode: small
num_classes:
width_mult: 1
---
Src/Super Gradients/Recipes/Arch Params/Pose Dekr Coco Rescoring Arch Params.Yaml (src/super_gradients/recipes/arch_params/pose_dekr_coco_rescoring_arch_params.yaml)
num_classes: 17
hidden_channels: 256
num_layers: 2
edge_links:
- [ 0, 1 ]
- [ 0, 2 ]
- [ 1, 2 ]
- [ 1, 3 ]
- [ 2, 4 ]
- [ 3, 5 ]
- [ 4, 6 ]
- [ 5, 6 ]
- [ 5, 7 ]
- [ 5, 11 ]
- [ 6, 8 ]
- [ 6, 12 ]
- [ 7, 9 ]
- [ 8, 10 ]
- [ 11, 12 ]
- [ 11, 13 ]
- [ 12, 14 ]
- [ 13, 15 ]
- [ 14, 16 ]
---
Src/Super Gradients/Recipes/Arch Params/Pose Dekr W32 No Dc Arch Params.Yaml (src/super_gradients/recipes/arch_params/pose_dekr_w32_no_dc_arch_params.yaml)
This model config is mimicing the one from the original repo:
https://github.com/HRNet/DEKR/blob/main/experiments/coco/w32/w32_4x_reg03_bs10_512_adam_lr1e-3_coco_x140.yaml
SPEC:
FINAL_CONV_KERNEL: 1
STAGES:
NUM_STAGES: 3
NUM_MODULES:
- 1
- 4
- 3
NUM_BRANCHES:
- 2
- 3
- 4
BLOCK:
- BASIC
- BASIC
- BASIC
NUM_BLOCKS:
- [4, 4]
- [4, 4, 4]
- [4, 4, 4, 4]
NUM_CHANNELS:
- [32, 64]
- [32, 64, 128]
- [32, 64, 128, 256]
FUSE_METHOD:
- SUM
- SUM
- SUM
HEAD_HEATMAP:
BLOCK: BASIC
NUM_BLOCKS: 1
NUM_CHANNELS: 32
DILATION_RATE: 1
HEATMAP_APPLY_SIGMOID: False
HEAD_OFFSET:
# Note we replace ADAPTIVE conv with BASIC conv since deformable conv is not supported in TensorRT,
# and we want the model to be exportable.
# Instead, we set dilation rate to 5 to mimic the effect of ADAPTIVE conv.
# Original recipe uses ADAPTIVE conv:
# BLOCK: ADAPTIVE
# DILATION_RATE: 1
BLOCK: BASIC # ADAPTIVE | BASIC
DILATION_RATE: 5
NUM_BLOCKS: 2
NUM_CHANNELS_PERKPT: 15
---
Src/Super Gradients/Recipes/Arch Params/Pose Pppose L Arch Params.Yaml (src/super_gradients/recipes/arch_params/pose_pppose_l_arch_params.yaml)
backbone:
CSPResNetBackbone:
layers: [ 3, 6, 6, 3 ] # Backbone's structure
channels: [ 64, 128, 256, 512, 1024 ] # Number of outputs channels for stem and consecutive feature maps
activation: silu
return_idx: [ 0, 1, 2, 3 ] # Indexes of feature maps to output, indiced 1,2,3 correspond to feature maps of stride 8,16,32
use_large_stem: True # If True, uses 3 conv+bn+act instead of 2 in stem blocks
use_alpha: False # If True, enables additional learnable weighting parameter for 1x1 branch in RepVGGBlock
depth_mult: 1
width_mult: 1
pretrained_weights:
neck:
CustomCSPPAN:
out_channels: [768, 384, 192, 128]
activation: silu
stage_num: 2
block_num: 2
spp: True
depth_mult: 1
width_mult: 1
heads:
LightweightDEKRHead:
heatmap_channels: 32
offset_channels_per_joint: 10
activation: silu
upscale_factor: 1
feature_map_index: -1
---
Src/Super Gradients/Recipes/Arch Params/Ppyoloe Arch Params.Yaml (src/super_gradients/recipes/arch_params/ppyoloe_arch_params.yaml)
depth_mult:
width_mult:
num_classes: 80
backbone:
layers: [ 3, 6, 6, 3 ] # Backbone's structure
channels: [ 64, 128, 256, 512, 1024 ] # Number of outputs channels for stem and consecutive feature maps
activation: silu
return_idx: [ 1, 2, 3 ] # Indexes of feature maps to output, indiced 1,2,3 correspond to feature maps of stride 8,16,32
use_large_stem: True # If True, uses 3 conv+bn+act instead of 2 in stem blocks
use_alpha: False # If True, enables additional learnable weighting parameter for 1x1 branch in RepVGGBlock
pretrained_weights:
neck:
in_channels: [256, 512, 1024]
out_channels: [768, 384, 192]
activation: silu
block_num: 3
stage_num: 1
spp: True
head:
in_channels: [768, 384, 192]
activation: silu
fpn_strides: [32, 16, 8]
grid_cell_scale: 5.0
grid_cell_offset: 0.5
reg_max: 16 # Number of bins for size prediction
eval_size: # Size of the image for evaluation. Setting this value can be beneficial for inference speed since anchors will not be regenerated for each forward call.
---
Src/Super Gradients/Recipes/Arch Params/Ppyoloe L Arch Params.Yaml (src/super_gradients/recipes/arch_params/ppyoloe_l_arch_params.yaml)
defaults:
- ppyoloe_arch_params
- _self_
depth_mult: 1.0
width_mult: 1.0
backbone:
pretrained_weights: https://deci-pretrained-models.s3.amazonaws.com/ppyolo_e/CSPResNetb_l_pretrained.pth
---
Src/Super Gradients/Recipes/Arch Params/Ppyoloe M Arch Params.Yaml (src/super_gradients/recipes/arch_params/ppyoloe_m_arch_params.yaml)
defaults:
- ppyoloe_arch_params
- _self_
depth_mult: 0.67
width_mult: 0.75
backbone:
pretrained_weights: https://deci-pretrained-models.s3.amazonaws.com/ppyolo_e/CSPResNetb_m_pretrained.pth
---
Src/Super Gradients/Recipes/Arch Params/Ppyoloe S Arch Params.Yaml (src/super_gradients/recipes/arch_params/ppyoloe_s_arch_params.yaml)
defaults:
- ppyoloe_arch_params
- _self_
depth_mult: 0.33
width_mult: 0.50
backbone:
pretrained_weights: https://deci-pretrained-models.s3.amazonaws.com/ppyolo_e/CSPResNetb_s_pretrained.pth
---
Src/Super Gradients/Recipes/Arch Params/Ppyoloe X Arch Params.Yaml (src/super_gradients/recipes/arch_params/ppyoloe_x_arch_params.yaml)
defaults:
- ppyoloe_arch_params
- _self_
depth_mult: 1.33
width_mult: 1.25
backbone:
pretrained_weights: https://deci-pretrained-models.s3.amazonaws.com/ppyolo_e/CSPResNetb_x_pretrained.pth
---
Src/Super Gradients/Recipes/Arch Params/RegnetY Arch Params.Yaml (src/super_gradients/recipes/arch_params/regnetY_arch_params.yaml)
backbone_mode: False # cuts off classification head
dropout_prob: 0 # dropout probability
droppath_prob: 0 # connection dropout probability
num_classes: 1000 # number of outputs of the classification head
---
Src/Super Gradients/Recipes/Arch Params/Repvgg Arch Params.Yaml (src/super_gradients/recipes/arch_params/repvgg_arch_params.yaml)
struct: # model's structure
num_classes: # number of outputs of the classification head
width_multiplier: # model's width multiplier
use_se: False # use squeeze and excitation layer
backbone_mode: False # cuts off classification head.
build_residual_branches: True # whether to add residual connections or not
in_channels: 3 # number of input channels
---
Src/Super Gradients/Recipes/Arch Params/Repvgga0 Arch Params.Yaml (src/super_gradients/recipes/arch_params/repvgga0_arch_params.yaml)
defaults:
- repvgg_arch_params
struct: [2, 4, 14, 1]
width_multiplier: [0.75, 0.75, 0.75, 2.5]
---
Src/Super Gradients/Recipes/Arch Params/Repvgga1 Arch Params.Yaml (src/super_gradients/recipes/arch_params/repvgga1_arch_params.yaml)
defaults:
- repvgg_arch_params
struct: [2, 4, 14, 1]
width_multiplier: [1, 1, 1, 2.5]
---
Src/Super Gradients/Recipes/Arch Params/Repvgga2 Arch Params.Yaml (src/super_gradients/recipes/arch_params/repvgga2_arch_params.yaml)
defaults:
- repvgg_arch_params
struct: [2, 4, 14, 1]
width_multiplier: [1.5, 1.5, 1.5, 2.75]
---
Src/Super Gradients/Recipes/Arch Params/Repvggb0 Arch Params.Yaml (src/super_gradients/recipes/arch_params/repvggb0_arch_params.yaml)
defaults:
- repvgg_arch_params
struct: [4, 6, 16, 1]
width_multiplier: [1, 1, 1, 2.5]
---
Src/Super Gradients/Recipes/Arch Params/Repvggb1 Arch Params.Yaml (src/super_gradients/recipes/arch_params/repvggb1_arch_params.yaml)
defaults:
- repvgg_arch_params
struct: [4, 6, 16, 1]
width_multiplier: [2, 2, 2, 4]
---
Src/Super Gradients/Recipes/Arch Params/Repvggb2 Arch Params.Yaml (src/super_gradients/recipes/arch_params/repvggb2_arch_params.yaml)
defaults:
- repvgg_arch_params
struct: [4, 6, 16, 1]
width_multiplier: [2.5, 2.5, 2.5, 5]
---
Src/Super Gradients/Recipes/Arch Params/Ssd Lite Mobilenetv2 Arch Params.Yaml (src/super_gradients/recipes/arch_params/ssd_lite_mobilenetv2_arch_params.yaml)
backbone:
MobileNetV2Backbone:
width_mult: 1.
structure:
grouped_conv_size: 1
out_layers: [['features', 14, 'conv', 2], ['features', 18]]
neck:
SSDInvertedResidualNeck:
blocks_out_channels: [512, 256, 256, 64]
expand_ratios: [0.2, 0.25, 0.5, 0.25]
grouped_conv_size: 1
heads:
SSDHead:
num_classes: 80
lite: True
anchors:
_target_: super_gradients.training.utils.ssd_utils.DefaultBoxes
fig_size: 320
feat_size: [20, 10, 5, 3, 2, 1]
scales: [32, 82, 133, 184, 235, 285, 336]
aspect_ratios: [[2, 3], [2, 3], [2, 3], [2, 3], [2, 3], [2, 3]]
scale_xy: 0.1
scale_wh: 0.2
---
Src/Super Gradients/Recipes/Arch Params/Ssd Mobilenetv1 Arch Params.Yaml (src/super_gradients/recipes/arch_params/ssd_mobilenetv1_arch_params.yaml)
backbone:
MobileNetV1Backbone:
out_layers: [['layers', 9]]
neck:
SSDBottleneckNeck:
blocks_out_channels: [1024, 512, 256, 256, 256]
bottleneck_channels: [256, 256, 128, 128, 128]
strides: [2, 2, 2, 1, 1]
kernel_sizes: [3, 3, 3, 3, 2]
heads:
SSDHead:
num_classes: 80
lite: False
anchors:
_target_: super_gradients.training.utils.ssd_utils.DefaultBoxes
fig_size: 320
feat_size: [40, 20, 10, 5, 3, 2]
scales: [22, 48, 106, 163, 221, 278, 336]
aspect_ratios: [[2], [2, 3], [2, 3], [2, 3], [2], [2]]
scale_xy: 0.1
scale_wh: 0.2
---
Src/Super Gradients/Recipes/Arch Params/Unet Arch Params.Yaml (src/super_gradients/recipes/arch_params/unet_arch_params.yaml)
defaults:
- unet_default_arch_params
- _self_
backbone_params:
strides_list: [1, 2, 2, 2, 2] # list of stride per stage.
width_list: [64, 128, 256, 512, 512] # list of num channels per stage.
num_blocks_list: [2, 2, 2, 2, 2] # list of num blocks per stage.
block_types_list: [ConvStage, ConvStage, ConvStage, ConvStage, ConvStage] # list of block types per stage. See unet_encoder.DownBlockType for options.
is_out_feature_list: [ True, True, True, True, True ] # list of flags whether stage features should be an output.
block_params:
downsample_mode: max_pool
context_module:
decoder_params:
# skip expansion ratio value, before fusing the skip features from the encoder with the decoder features, a projection
# convolution is applied upon the encoder features to project the num_channels by skip_expansion.
skip_expansion: 1.
decoder_scale: .5 # num_channels width ratio between encoder stages and decoder stages.
up_block_types: [UpCatBlock, UpCatBlock, UpCatBlock, UpCatBlock] # See unet_decoder.UpBlockType for options.
up_block_repeat_list: [ 2, 2, 2, 1] # num of blocks per decoder stage, the block implementation depends on the up-block type.
mode: bilinear
align_corners: False
up_factor: 2
final_upsample_factor: 1 # Final upsample scale factor after the segmentation head.
_convert_: all
---
Src/Super Gradients/Recipes/Arch Params/Unet Default Arch Params.Yaml (src/super_gradients/recipes/arch_params/unet_default_arch_params.yaml)
backbone_params:
in_channels: 3
strides_list: [2, 2, 2, 2, 2] # list of stride per stage.
width_list: [32, 64, 128, 256, 512] # list of num channels per stage.
num_blocks_list: [2, 3, 3, 2, 2] # list of num blocks per stage.
block_types_list: [RepVGGStage, RepVGGStage, RepVGGStage, RepVGGStage, RepVGGStage] # list of block types per stage. See unet_encoder.DownBlockType for options.
is_out_feature_list: [ True, True, True, True, True ] # list of flags whether stage features should be an output.
block_params:
downsample_mode: anti_alias # RepVGG stage param
droppath_prob: 0. # XBlock stage param
bottleneck_ratio: 1. # XBlock stage param
group_width: 16 # XBlock stage param
se_ratio: # XBlock stage param
steps: 4 # STDC stage params
stdc_downsample_mode: dw_conv # STDC stage params
context_module:
ASPP:
in_channels: ${last:${arch_params.backbone_params.width_list}}
dilation_list: [2, 4, 6]
in_out_ratio: 1.
# legacy parameter to support old trained checkpoints that were trained by mistake with extra redundant
# biases before batchnorm operators. should be set to False for new training processes.
use_bias: False
decoder_params:
# skip expansion ratio value, before fusing the skip features from the encoder with the decoder features, a projection
# convolution is applied upon the encoder features to project the num_channels by skip_expansion.
skip_expansion: 0.25
decoder_scale: 0.25 # num_channels width ratio between encoder stages and decoder stages.
up_block_types: [UpCatBlock, UpCatBlock, UpCatBlock, UpCatBlock] # See unet_decoder.UpBlockType for options.
up_block_repeat_list: [ 1, 1, 1, 1] # num of blocks per decoder stage, the block implementation depends on the up-block type.
mode: bilinear
fallback_mode:
align_corners: False
up_factor: 2
is_skip_list: [True, True, True, True] # List of flags whether to use feature-map from encoder stage as skip connection or not.
min_decoder_channels: 1 # The minimum num_channels of decoder stages. Useful i.e if we want to keep the width above the num of classes.
dropout: 0.
final_upsample_factor: 2 # Final upsample scale factor after the segmentation head.
head_upsample_mode: bilinear
align_corners: False
head_hidden_channels: # num channels before the last classification layer. see mid_channels in SegmentationHead class.
use_aux_heads: False
aux_heads_params:
use_aux_list: [False, False, True, True, True] # whether to append to auxiliary head per encoder stage.
aux_heads_factor: [2, 4, 8, 16, 32] # Upsample factor per encoder stage.
aux_hidden_channels: [32, 32, 64, 64, 64] # Hidden num channels before last classification layer, per encoder stage.
aux_out_channels: [1, 1, 19, 19, 19] # Output channels, can be refers as num_classes, of auxiliary head per encoder stage.
_convert_: all
---
Src/Super Gradients/Recipes/Arch Params/Vit Base Arch Params.Yaml (src/super_gradients/recipes/arch_params/vit_base_arch_params.yaml)
num_classes: 1000
image_size: [224, 224]
patch_size: [16, 16]
---
Src/Super Gradients/Recipes/Arch Params/Yolo Arch Params.Yaml (src/super_gradients/recipes/arch_params/yolo_arch_params.yaml)
anchors:
_target_: super_gradients.training.utils.detection_utils.Anchors
anchors_list: [[10, 13, 16, 30, 33, 23],[30, 61, 62, 45, 59, 119],[116, 90, 156, 198, 373, 326]]
strides: [8, 16, 32]
num_classes: 80 # Number of classes to predict
depth_mult_factor: 1.0 # depth multiplier for the entire model, overridden for predefined YoloV5S, YoloV5M, YoloV5L
width_mult_factor: 1.0 # width multiplier for the entire model, overridden for predefined YoloV5S, YoloV5M, YoloV5L
channels_in: 3 # Number of channels in the input image
skip_connections_list: [[12, [6]], [16, [4]], [19, [14]], [22, [10]], [24, [17, 20]]]
A list defining skip connections. format is [target: [source1, source2, ...]]. Each item defines a skip
connection from all sources to the target according to the layers index (count starts from the backbone)
backbone_connection_channels: [1024, 512, 256] # width of backbone channels that are concatenated with the head
scaled_backbone_width: True # True if width_mult_factor is applied to the backbone
(is the case with the default backbones)
which means that backbone_connection_channels should be used with a width_mult_factor
False if backbone_connection_channels should be used as is
fuse_conv_and_bn: False # Fuse sequential Conv + B.N layers into a single one
add_nms: False # Add the NMS module to the computational graph
nms_conf: 0.25 # When add_nms is True during NMS predictions with confidence lower than this will be discarded
nms_iou: 0.45 # When add_nms is True IoU threshold for NMS algorithm
# (with smaller value more boxed will be considered "the same" and removed)
yolo_type: 'yolox' # Type of yolo to build: 'yoloX' is th only type currently supported.
stem_type: # 'focus' and '6x6' are supported, by default is defined by yolo_type and yolo_version
depthwise: False # use depthwise separable convolutions all over the model
xhead_inter_channels: # (has an impact only if yolo_type is yoloX)
Channels in classification and regression branches of the detecting blocks
if is None the first of input channels will be used by default
xhead_groups: # (has an impact only if yolo_type is yoloX)
Num groups in convs in classification and regression branches of the detecting blocks
if None default groups will be used according to conv type
(1 for Conv and depthwise for GroupedConvBlock)
_convert_: all
---
Src/Super Gradients/Recipes/Arch Params/Yolo Nas L Arch Params.Yaml (src/super_gradients/recipes/arch_params/yolo_nas_l_arch_params.yaml)
in_channels: 3
backbone:
NStageBackbone:
stem:
YoloNASStem:
out_channels: 48
stages:
- YoloNASStage:
out_channels: 96
num_blocks: 2
activation_type: relu
hidden_channels: 96
concat_intermediates: True
- YoloNASStage:
out_channels: 192
num_blocks: 3
activation_type: relu
hidden_channels: 128
concat_intermediates: True
- YoloNASStage:
out_channels: 384
num_blocks: 5
activation_type: relu
hidden_channels: 256
concat_intermediates: True
- YoloNASStage:
out_channels: 768
num_blocks: 2
activation_type: relu
hidden_channels: 512
concat_intermediates: True
context_module:
SPP:
output_channels: 768
activation_type: relu
k: [5,9,13]
out_layers: [stage1, stage2, stage3, context_module]
neck:
YoloNASPANNeckWithC2:
neck1:
YoloNASUpStage:
out_channels: 192
num_blocks: 4
hidden_channels: 128
width_mult: 1
depth_mult: 1
activation_type: relu
reduce_channels: True
neck2:
YoloNASUpStage:
out_channels: 96
num_blocks: 4
hidden_channels: 128
width_mult: 1
depth_mult: 1
activation_type: relu
reduce_channels: True
neck3:
YoloNASDownStage:
out_channels: 192
num_blocks: 4
hidden_channels: 128
activation_type: relu
width_mult: 1
depth_mult: 1
neck4:
YoloNASDownStage:
out_channels: 384
num_blocks: 4
hidden_channels: 256
activation_type: relu
width_mult: 1
depth_mult: 1
heads:
NDFLHeads:
num_classes: 80
reg_max: 16
heads_list:
- YoloNASDFLHead:
inter_channels: 128
width_mult: 1
first_conv_group_size: 0
stride: 8
- YoloNASDFLHead:
inter_channels: 256
width_mult: 1
first_conv_group_size: 0
stride: 16
- YoloNASDFLHead:
inter_channels: 512
width_mult: 1
first_conv_group_size: 0
stride: 32
bn_eps: 1e-3
bn_momentum: 0.03
inplace_act: True
_convert_: all
---
Src/Super Gradients/Recipes/Arch Params/Yolo Nas M Arch Params.Yaml (src/super_gradients/recipes/arch_params/yolo_nas_m_arch_params.yaml)
in_channels: 3
backbone:
NStageBackbone:
stem:
YoloNASStem:
out_channels: 48
stages:
- YoloNASStage:
out_channels: 96
num_blocks: 2
activation_type: relu
hidden_channels: 64
concat_intermediates: True
- YoloNASStage:
out_channels: 192
num_blocks: 3
activation_type: relu
hidden_channels: 128
concat_intermediates: True
- YoloNASStage:
out_channels: 384
num_blocks: 5
activation_type: relu
hidden_channels: 256
concat_intermediates: True
- YoloNASStage:
out_channels: 768
num_blocks: 2
activation_type: relu
hidden_channels: 384
concat_intermediates: False
context_module:
SPP:
output_channels: 768
activation_type: relu
k: [5,9,13]
out_layers: [stage1, stage2, stage3, context_module]
neck:
YoloNASPANNeckWithC2:
neck1:
YoloNASUpStage:
out_channels: 192
num_blocks: 2
hidden_channels: 192
width_mult: 1
depth_mult: 1
activation_type: relu
reduce_channels: True
neck2:
YoloNASUpStage:
out_channels: 96
num_blocks: 3
hidden_channels: 64
width_mult: 1
depth_mult: 1
activation_type: relu
reduce_channels: True
neck3:
YoloNASDownStage:
out_channels: 192
num_blocks: 2
hidden_channels: 192
activation_type: relu
width_mult: 1
depth_mult: 1
neck4:
YoloNASDownStage:
out_channels: 384
num_blocks: 3
hidden_channels: 256
activation_type: relu
width_mult: 1
depth_mult: 1
heads:
NDFLHeads:
num_classes: 80
reg_max: 16
heads_list:
- YoloNASDFLHead:
inter_channels: 128
width_mult: 0.75
first_conv_group_size: 0
stride: 8
- YoloNASDFLHead:
inter_channels: 256
width_mult: 0.75
first_conv_group_size: 0
stride: 16
- YoloNASDFLHead:
inter_channels: 512
width_mult: 0.75
first_conv_group_size: 0
stride: 32
bn_eps: 1e-3
bn_momentum: 0.03
inplace_act: True
_convert_: all
---
Src/Super Gradients/Recipes/Arch Params/Yolo Nas Pose L Arch Params.Yaml (src/super_gradients/recipes/arch_params/yolo_nas_pose_l_arch_params.yaml)
in_channels: 3
backbone:
NStageBackbone:
stem:
YoloNASStem:
out_channels: 48
stages:
- YoloNASStage:
out_channels: 96
num_blocks: 2
activation_type: relu
hidden_channels: 96
concat_intermediates: True
- YoloNASStage:
out_channels: 192
num_blocks: 3
activation_type: relu
hidden_channels: 128
concat_intermediates: True
- YoloNASStage:
out_channels: 384
num_blocks: 5
activation_type: relu
hidden_channels: 256
concat_intermediates: True
- YoloNASStage:
out_channels: 768
num_blocks: 2
activation_type: relu
hidden_channels: 512
concat_intermediates: True
context_module:
SPP:
output_channels: 768
activation_type: relu
k: [5,9,13]
out_layers: [stage1, stage2, stage3, context_module]
neck:
YoloNASPANNeckWithC2:
neck1:
YoloNASUpStage:
out_channels: 192
num_blocks: 4
hidden_channels: 128
width_mult: 1
depth_mult: 1
activation_type: relu
reduce_channels: True
neck2:
YoloNASUpStage:
out_channels: 96
num_blocks: 4
hidden_channels: 128
width_mult: 1
depth_mult: 1
activation_type: relu
reduce_channels: True
neck3:
YoloNASDownStage:
out_channels: 192
num_blocks: 4
hidden_channels: 128
activation_type: relu
width_mult: 1
depth_mult: 1
neck4:
YoloNASDownStage:
out_channels: 384
num_blocks: 4
hidden_channels: 256
activation_type: relu
width_mult: 1
depth_mult: 1
heads:
YoloNASPoseNDFLHeads:
num_classes: 17
reg_max: 16
heads_list:
- YoloNASPoseDFLHead:
bbox_inter_channels: 128
pose_inter_channels: 128
pose_regression_blocks: 2
shared_stem: False
width_mult: 1
pose_conf_in_class_head: True
pose_block_use_repvgg: False
first_conv_group_size: 0
num_classes:
stride: 8
reg_max: 16
cls_dropout_rate: 0.0
reg_dropout_rate: 0.0
- YoloNASPoseDFLHead:
bbox_inter_channels: 256
pose_inter_channels: 512
pose_regression_blocks: 2
shared_stem: False
width_mult: 1
pose_conf_in_class_head: True
pose_block_use_repvgg: False
first_conv_group_size: 0
num_classes:
stride: 16
reg_max: 16
cls_dropout_rate: 0.0
reg_dropout_rate: 0.0
- YoloNASPoseDFLHead:
bbox_inter_channels: 512
pose_inter_channels: 512
pose_regression_blocks: 3
shared_stem: False
width_mult: 1
pose_conf_in_class_head: True
pose_block_use_repvgg: False
first_conv_group_size: 0
num_classes:
stride: 32
reg_max: 16
cls_dropout_rate: 0.0
reg_dropout_rate: 0.0
bn_eps: 1e-6
bn_momentum: 0.03
inplace_act: True
_convert_: all
---
Src/Super Gradients/Recipes/Arch Params/Yolo Nas Pose M Arch Params.Yaml (src/super_gradients/recipes/arch_params/yolo_nas_pose_m_arch_params.yaml)
in_channels: 3
backbone:
NStageBackbone:
stem:
YoloNASStem:
out_channels: 48
stages:
- YoloNASStage:
out_channels: 96
num_blocks: 2
activation_type: relu
hidden_channels: 64
concat_intermediates: True
- YoloNASStage:
out_channels: 192
num_blocks: 3
activation_type: relu
hidden_channels: 128
concat_intermediates: True
- YoloNASStage:
out_channels: 384
num_blocks: 5
activation_type: relu
hidden_channels: 256
concat_intermediates: True
- YoloNASStage:
out_channels: 768
num_blocks: 2
activation_type: relu
hidden_channels: 384
concat_intermediates: False
context_module:
SPP:
output_channels: 768
activation_type: relu
k: [5,9,13]
out_layers: [stage1, stage2, stage3, context_module]
neck:
YoloNASPANNeckWithC2:
neck1:
YoloNASUpStage:
out_channels: 192
num_blocks: 2
hidden_channels: 192
width_mult: 1
depth_mult: 1
activation_type: relu
reduce_channels: True
neck2:
YoloNASUpStage:
out_channels: 96
num_blocks: 3
hidden_channels: 64
width_mult: 1
depth_mult: 1
activation_type: relu
reduce_channels: True
neck3:
YoloNASDownStage:
out_channels: 192
num_blocks: 2
hidden_channels: 192
activation_type: relu
width_mult: 1
depth_mult: 1
neck4:
YoloNASDownStage:
out_channels: 384
num_blocks: 3
hidden_channels: 256
activation_type: relu
width_mult: 1
depth_mult: 1
heads:
YoloNASPoseNDFLHeads:
num_classes: 17
reg_max: 16
pose_offset_multiplier: 1.0
compensate_grid_cell_offset: True
inference_mode: False # True used only when benchmarking
heads_list:
- YoloNASPoseDFLHead:
bbox_inter_channels: 128
pose_inter_channels: 128
pose_regression_blocks: 2
shared_stem: False
width_mult: 0.75
pose_conf_in_class_head: True
pose_block_use_repvgg: False
first_conv_group_size: 0
num_classes:
stride: 8
reg_max: 16
- YoloNASPoseDFLHead:
bbox_inter_channels: 256
pose_inter_channels: 512
pose_regression_blocks: 2
shared_stem: False
width_mult: 0.75
pose_conf_in_class_head: True
pose_block_use_repvgg: False
first_conv_group_size: 0
num_classes:
stride: 16
reg_max: 16
- YoloNASPoseDFLHead:
bbox_inter_channels: 512
pose_inter_channels: 512
pose_regression_blocks: 3
shared_stem: False
width_mult: 0.75
pose_conf_in_class_head: True
pose_block_use_repvgg: False
first_conv_group_size: 0
num_classes:
stride: 32
reg_max: 16
bn_eps: 1e-6
bn_momentum: 0.1
inplace_act: True
_convert_: all
---
Src/Super Gradients/Recipes/Arch Params/Yolo Nas Pose N Arch Params.Yaml (src/super_gradients/recipes/arch_params/yolo_nas_pose_n_arch_params.yaml)
in_channels: 3
backbone:
NStageBackbone:
stem:
YoloNASStem:
out_channels: 32
stages:
- YoloNASStage:
out_channels: 64
num_blocks: 2
activation_type: relu
hidden_channels: 32
concat_intermediates: False
- YoloNASStage:
out_channels: 128
num_blocks: 3
activation_type: relu
hidden_channels: 48
concat_intermediates: False
- YoloNASStage:
out_channels: 256
num_blocks: 4
activation_type: relu
hidden_channels: 64
concat_intermediates: False
- YoloNASStage:
out_channels: 512
num_blocks: 2
activation_type: relu
hidden_channels: 128
concat_intermediates: False
context_module:
SPP:
output_channels: 512
activation_type: relu
k: [5,9,13]
out_layers: [stage1, stage2, stage3, context_module]
neck:
YoloNASPANNeckWithC2:
neck1:
YoloNASUpStage:
out_channels: 128
num_blocks: 2
hidden_channels: 48
width_mult: 1
depth_mult: 1
activation_type: relu
reduce_channels: True
neck2:
YoloNASUpStage:
out_channels: 64
num_blocks: 2
hidden_channels: 32
width_mult: 1
depth_mult: 1
activation_type: relu
reduce_channels: True
neck3:
YoloNASDownStage:
out_channels: 128
num_blocks: 2
hidden_channels: 48
activation_type: relu
width_mult: 1
depth_mult: 1
neck4:
YoloNASDownStage:
out_channels: 256
num_blocks: 2
hidden_channels: 48
activation_type: relu
width_mult: 1
depth_mult: 1
heads:
YoloNASPoseNDFLHeads:
num_classes: 17
reg_max: 16
pose_offset_multiplier: 1.0
compensate_grid_cell_offset: True
inference_mode: False # True used only when benchmarking
heads_list:
- YoloNASPoseDFLHead:
bbox_inter_channels: 128
pose_inter_channels: 128
pose_regression_blocks: 2
shared_stem: False
width_mult: 0.33
pose_conf_in_class_head: True
pose_block_use_repvgg: False
first_conv_group_size: 0
num_classes:
stride: 8
reg_max: 16
- YoloNASPoseDFLHead:
bbox_inter_channels: 256
pose_inter_channels: 512
pose_regression_blocks: 2
shared_stem: False
width_mult: 0.33
pose_conf_in_class_head: True
pose_block_use_repvgg: False
first_conv_group_size: 0
num_classes:
stride: 16
reg_max: 16
- YoloNASPoseDFLHead:
bbox_inter_channels: 512
pose_inter_channels: 512
pose_regression_blocks: 3
shared_stem: False
width_mult: 0.33
pose_conf_in_class_head: True
pose_block_use_repvgg: False
first_conv_group_size: 0
num_classes:
stride: 32
reg_max: 16
bn_eps: 1e-6
bn_momentum: 0.03
inplace_act: True
_convert_: all
---
Src/Super Gradients/Recipes/Arch Params/Yolo Nas Pose S Arch Params.Yaml (src/super_gradients/recipes/arch_params/yolo_nas_pose_s_arch_params.yaml)
in_channels: 3
backbone:
NStageBackbone:
stem:
YoloNASStem:
out_channels: 48
stages:
- YoloNASStage:
out_channels: 96
num_blocks: 2
activation_type: relu
hidden_channels: 32
concat_intermediates: False
- YoloNASStage:
out_channels: 192
num_blocks: 3
activation_type: relu
hidden_channels: 64
concat_intermediates: False
- YoloNASStage:
out_channels: 384
num_blocks: 5
activation_type: relu
hidden_channels: 96
concat_intermediates: False
- YoloNASStage:
out_channels: 768
num_blocks: 2
activation_type: relu
hidden_channels: 192
concat_intermediates: False
context_module:
SPP:
output_channels: 768
activation_type: relu
k: [5,9,13]
out_layers: [stage1, stage2, stage3, context_module]
neck:
YoloNASPANNeckWithC2:
neck1:
YoloNASUpStage:
out_channels: 192
num_blocks: 2
hidden_channels: 64
width_mult: 1
depth_mult: 1
activation_type: relu
reduce_channels: True
neck2:
YoloNASUpStage:
out_channels: 96
num_blocks: 2
hidden_channels: 48
width_mult: 1
depth_mult: 1
activation_type: relu
reduce_channels: True
neck3:
YoloNASDownStage:
out_channels: 192
num_blocks: 2
hidden_channels: 64
activation_type: relu
width_mult: 1
depth_mult: 1
neck4:
YoloNASDownStage:
out_channels: 384
num_blocks: 2
hidden_channels: 64
activation_type: relu
width_mult: 1
depth_mult: 1
heads:
YoloNASPoseNDFLHeads:
num_classes: 17
reg_max: 16
pose_offset_multiplier: 1.0
compensate_grid_cell_offset: True
inference_mode: False # True used only when benchmarking
heads_list:
- YoloNASPoseDFLHead:
bbox_inter_channels: 128
pose_inter_channels: 128
pose_regression_blocks: 2
shared_stem: False
width_mult: 0.5
pose_conf_in_class_head: True
pose_block_use_repvgg: False
first_conv_group_size: 0
num_classes:
stride: 8
reg_max: 16
- YoloNASPoseDFLHead:
bbox_inter_channels: 256
pose_inter_channels: 512
pose_regression_blocks: 2
shared_stem: False
width_mult: 0.5
pose_conf_in_class_head: True
pose_block_use_repvgg: False
first_conv_group_size: 0
num_classes:
stride: 16
reg_max: 16
- YoloNASPoseDFLHead:
bbox_inter_channels: 512
pose_inter_channels: 512
pose_regression_blocks: 3
shared_stem: False
width_mult: 0.5
pose_conf_in_class_head: True
pose_block_use_repvgg: False
first_conv_group_size: 0
num_classes:
stride: 32
reg_max: 16
bn_eps: 1e-6
bn_momentum: 0.1
inplace_act: True
_convert_: all
---
Src/Super Gradients/Recipes/Arch Params/Yolo Nas S Arch Params.Yaml (src/super_gradients/recipes/arch_params/yolo_nas_s_arch_params.yaml)
in_channels: 3
backbone:
NStageBackbone:
stem:
YoloNASStem:
out_channels: 48
stages:
- YoloNASStage:
out_channels: 96
num_blocks: 2
activation_type: relu
hidden_channels: 32
concat_intermediates: False
- YoloNASStage:
out_channels: 192
num_blocks: 3
activation_type: relu
hidden_channels: 64
concat_intermediates: False
- YoloNASStage:
out_channels: 384
num_blocks: 5
activation_type: relu
hidden_channels: 96
concat_intermediates: False
- YoloNASStage:
out_channels: 768
num_blocks: 2
activation_type: relu
hidden_channels: 192
concat_intermediates: False
context_module:
SPP:
output_channels: 768
activation_type: relu
k: [5,9,13]
out_layers: [stage1, stage2, stage3, context_module]
neck:
YoloNASPANNeckWithC2:
neck1:
YoloNASUpStage:
out_channels: 192
num_blocks: 2
hidden_channels: 64
width_mult: 1
depth_mult: 1
activation_type: relu
reduce_channels: True
neck2:
YoloNASUpStage:
out_channels: 96
num_blocks: 2
hidden_channels: 48
width_mult: 1
depth_mult: 1
activation_type: relu
reduce_channels: True
neck3:
YoloNASDownStage:
out_channels: 192
num_blocks: 2
hidden_channels: 64
activation_type: relu
width_mult: 1
depth_mult: 1
neck4:
YoloNASDownStage:
out_channels: 384
num_blocks: 2
hidden_channels: 64
activation_type: relu
width_mult: 1
depth_mult: 1
heads:
NDFLHeads:
num_classes: 80
reg_max: 16
heads_list:
- YoloNASDFLHead:
inter_channels: 128
width_mult: 0.5
first_conv_group_size: 0
stride: 8
- YoloNASDFLHead:
inter_channels: 256
width_mult: 0.5
first_conv_group_size: 0
stride: 16
- YoloNASDFLHead:
inter_channels: 512
width_mult: 0.5
first_conv_group_size: 0
stride: 32
bn_eps: 1e-3
bn_momentum: 0.03
inplace_act: True
_convert_: all
---
Src/Super Gradients/Recipes/Arch Params/Yolox L Arch Params.Yaml (src/super_gradients/recipes/arch_params/yolox_l_arch_params.yaml)
defaults:
- yolo_arch_params
anchors:
_target_: super_gradients.training.utils.detection_utils.Anchors
anchors_list: [[0,0], [0,0], [0,0]]
strides: [8, 16, 32]
yolo_type: 'yoloX'
depth_mult_factor: 1.0
width_mult_factor: 1.0
---
Src/Super Gradients/Recipes/Arch Params/Yolox M Arch Params.Yaml (src/super_gradients/recipes/arch_params/yolox_m_arch_params.yaml)
defaults:
- yolo_arch_params
anchors:
_target_: super_gradients.training.utils.detection_utils.Anchors
anchors_list: [[0,0], [0,0], [0,0]]
strides: [8, 16, 32]
yolo_type: 'yoloX'
depth_mult_factor: 0.67
width_mult_factor: 0.75
---
Src/Super Gradients/Recipes/Arch Params/Yolox Nano Arch Params.Yaml (src/super_gradients/recipes/arch_params/yolox_nano_arch_params.yaml)
defaults:
- yolo_arch_params
anchors:
_target_: super_gradients.training.utils.detection_utils.Anchors
anchors_list: [[0,0], [0,0], [0,0]]
strides: [8, 16, 32]
yolo_type: 'yoloX'
depth_mult_factor: 0.33
width_mult_factor: 0.25
---
Src/Super Gradients/Recipes/Arch Params/Yolox S Arch Params.Yaml (src/super_gradients/recipes/arch_params/yolox_s_arch_params.yaml)
defaults:
- yolo_arch_params
anchors:
_target_: super_gradients.training.utils.detection_utils.Anchors
anchors_list: [[0,0], [0,0], [0,0]]
strides: [8, 16, 32]
yolo_type: 'yoloX'
depth_mult_factor: 0.33
width_mult_factor: 0.5
---
Src/Super Gradients/Recipes/Arch Params/Yolox Tiny Arch Params.Yaml (src/super_gradients/recipes/arch_params/yolox_tiny_arch_params.yaml)
defaults:
- yolo_arch_params
anchors:
_target_: super_gradients.training.utils.detection_utils.Anchors
anchors_list: [[], [], []]
strides: [8, 16, 32]
yolo_type: 'yoloX'
depth_mult_factor: 0.33
width_mult_factor: 0.375
---
Src/Super Gradients/Recipes/Arch Params/Yolox X Arch Params.Yaml (src/super_gradients/recipes/arch_params/yolox_x_arch_params.yaml)
defaults:
- yolo_arch_params
anchors:
_target_: super_gradients.training.utils.detection_utils.Anchors
anchors_list: [[], [], []]
strides: [8, 16, 32]
yolo_type: 'yoloX'
depth_mult_factor: 1.33
width_mult_factor: 1.25
---
Src/Super Gradients/Recipes/Checkpoint Params/Default Checkpoint Params.Yaml (src/super_gradients/recipes/checkpoint_params/default_checkpoint_params.yaml)
load_checkpoint: False # whether to load checkpoint
load_backbone: False # whether to load only backbone part of checkpoint
checkpoint_path: # checkpoint path that is located in super_gradients/checkpoints
external_checkpoint_path: # checkpoint path that is not located in super_gradients/checkpoints
source_ckpt_folder_name: # dirname for checkpoint loading
strict_load: # key matching strictness for loading checkpoint's weights
_target_: super_gradients.training.sg_trainer.StrictLoad
value: no_key_matching
pretrained_weights: # a string describing the dataset of the pretrained weights (for example "imagenent").
num_classes of checkpoint_path/ pretrained_weights, when checkpoint_path is not None.
Used when num_classes != checkpoint_num_class.
In this case, the module will be initialized with checkpoint_num_class, then weights will be loaded.
Finally model.replace_head(new_num_classes=num_classes) is called to replace the head with new_num_classes.
checkpoint_num_classes: # number of classes in the checkpoint
---
Src/Super Gradients/Recipes/Checkpoint Params/Vit Base Imagenet Checkpoint Params.Yaml (src/super_gradients/recipes/checkpoint_params/vit_base_imagenet_checkpoint_params.yaml)
defaults:
- default_checkpoint_params
pretrained_weights: imagenet21k
---
Src/Super Gradients/Recipes/Conversion Params/Cifar10 Conversion Params.Yaml (src/super_gradients/recipes/conversion_params/cifar10_conversion_params.yaml)
Example conversion parameters, to be used with super_gradients/examples/convert_recipe_example/convert_recipe_example.py
Suppose you trained cifar10_resnet using train_from_recipe beforehand, Then:
python convert_recipe_example.py --config-name=cifar10_conversion_params experiment_name=YOUR_EXPERIMENT_NAME.
Alternatively (or if ckpts are located anywhere else from the default checkpoints dir), you can give the full checkpoint path:
python convert_recipe_example.py --config-name=cifar10_conversion_params checkpoint_path=YOUR_CHECKPOINT_PATH
defaults:
- default_conversion_params
- _self_
experiment_name: resnet18_cifar # The experiment name used to train the model (optional- ignored when checkpoint_path is given)
CONVERSION RELATED PARAMS
out_path: # str, Destination path for the .onnx file. When None- out_path will be the resolved checkpoint path replacing .ckpt suffix with .onnx.
input_shape: # input shape, not including batch_size. Always channels first (i.e (3, 224, 224)).
- 3
- 32
- 32
pre_process: # Preprocessing pipeline, will be resolved by TransformsFactory(), and will be baked into the converted model (optional).
Compose:
transforms:
- Standardize
- Normalize:
mean:
- 0.4914
- 0.4822
- 0.4465
std:
- 0.2023
- 0.1994
- 0.2010
post_process: # Postprocessing pipeline, will be resolved by TransformsFactory(), and will be baked into the converted model (optional).
prep_model_for_conversion_kwargs: # For SgModules, args to be passed to model.prep_model_for_conversion prior to torch.onnx.export call.
torch_onnx_export_kwargs: # kwargs (EXCLUDING: FIRST 3 KWARGS- MODEL, F, ARGS). to be unpacked in torch.onnx.export call
opset_version: 16
---
Src/Super Gradients/Recipes/Conversion Params/Default Conversion Params.Yaml (src/super_gradients/recipes/conversion_params/default_conversion_params.yaml)
experiment_name: # The experiment name used to train the model (optional- ignored when checkpoint_path is given)
run_id: # The directory name of the required checkpoint i.e. RUN_20230823_154026_757034 - if left empty, the last run will be used
ckpt_root_dir: # The checkpoint root directory, s.t ckpt_root_dir/experiment_name/ckpt_name resides.
# Can be ignored if the checkpoints directory is the default (i.e path to checkpoints module from contents root), or when checkpoint_path is given
ckpt_name: ckpt_best.pth # Name of the checkpoint to export ("ckpt_latest.pth", "average_model.pth" or "ckpt_best.pth" for instance).
checkpoint_path:
strict_load: no_key_matching # One of [On, Off, no_key_matching] (case insensitive) See super_gradients/common/data_types/enum/strict_load.py
NOTES ON: ckpt_root_dir, checkpoint_path, and ckpt_name:
- ckpt_root_dir, experiment_name and ckpt_name are only used when checkpoint_path is None.
- when checkpoint_path is None, the model will be vuilt according to the output yaml config inside ckpt_root_dir/experiment_name/ckpt_name. Also note that in
this case its also legal not to pass ckpt_root_dir, which will be resolved to the default SG ckpt dir.
CONVERSION RELATED PARAMS
out_path: # str, Destination path for the .onnx file. When None- will be set to the checkpoint_path.replace(".ckpt",".onnx").
input_shape: # DEPRECATED USE input_size KWARG IN prep_model_for_conversion_kwargs INSTEAD.
pre_process: # Preprocessing pipeline, will be resolved by TransformsFactory(), and will be baked into the converted model (optional).
post_process: # Postprocessing pipeline, will be resolved by TransformsFactory(), and will be baked into the converted model (optional).
prep_model_for_conversion_kwargs: # For SgModules, args to be passed to model.prep_model_for_conversion prior to torch.onnx.export call.
torch_onnx_export_kwargs: # kwargs (EXCLUDING: FIRST 3 KWARGS- MODEL, F, ARGS). to be unpacked in torch.onnx.export call
simplify: True # whether to apply onnx simplifier method, same as python -m onnxsim onnx_path onnx_sim_path. When true, the simplified models will be saved in out_path.
---
Src/Super Gradients/Recipes/Dataset Params/Cifar100 Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/cifar100_dataset_params.yaml)
train_dataset_params:
root: /data/cifar100
train: True
transforms:
- RandomCrop:
size: 32
padding: 4
- RandomHorizontalFlip
- ToTensor
- Normalize:
mean:
- 0.5071
- 0.4865
- 0.4409
std:
- 0.2673
- 0.2564
- 0.2762
target_transform: null
download: True
train_dataloader_params:
shuffle: True
batch_size: 256
num_workers: 8
drop_last: False
pin_memory: True
val_dataset_params:
root: /data/cifar100
train: False
transforms:
- Resize:
size: 32
- ToTensor
- Normalize:
mean:
- 0.5071
- 0.4865
- 0.4409
std:
- 0.2673
- 0.2564
- 0.2762
target_transform: null
download: True
val_dataloader_params:
batch_size: 512
num_workers: 8
drop_last: False
pin_memory: True
---
Src/Super Gradients/Recipes/Dataset Params/Cifar10 Albumentations Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/cifar10_albumentations_dataset_params.yaml)
Equivalent to cifar10_dataset_params.yaml, but uses albumentations transforms.
The purpose of the below configuration is to demonstrate the use of Albumentation transforms in train_from_recipe.
batch_size: 256 # batch size for trainset
val_batch_size: 512 # batch size for valset in DatasetInterface
TODO: REMOVE ABOVE, HERE FOR COMPATIBILITY UNTIL WE REMOVE DATASET_INTERFACE
train_dataset_params:
root: ./data/cifar10
train: True
transforms:
Albumentations:
Compose:
transforms:
- RandomCrop:
height: 32
width: 32
- HorizontalFlip:
p: 0.5
- Normalize:
mean:
- 0.4914
- 0.4822
- 0.4465
std:
- 0.2023
- 0.1994
- 0.2010
- ToTensorV2
target_transform: null
download: True
train_dataloader_params:
shuffle: True
batch_size: 256
num_workers: 8
drop_last: False
pin_memory: True
val_dataset_params:
root: ./data/cifar10
train: False
transforms:
Albumentations:
Compose:
transforms:
- Normalize:
mean:
- 0.4914
- 0.4822
- 0.4465
std:
- 0.2023
- 0.1994
- 0.2010
- ToTensorV2
target_transform: null
download: True
val_dataloader_params:
batch_size: 512
num_workers: 8
drop_last: False
pin_memory: True
---
Src/Super Gradients/Recipes/Dataset Params/Cifar10 Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/cifar10_dataset_params.yaml)
batch_size: 256 # batch size for trainset
val_batch_size: 512 # batch size for valset in DatasetInterface
TODO: REMOVE ABOVE, HERE FOR COMPATIBILITY UNTIL WE REMOVE DATASET_INTERFACE
train_dataset_params:
root: ./data/cifar10
train: True
transforms:
- RandomCrop:
size: 32
padding: 4
- RandomHorizontalFlip
- ToTensor
- Normalize:
mean:
- 0.4914
- 0.4822
- 0.4465
std:
- 0.2023
- 0.1994
- 0.2010
target_transform: null
download: True
train_dataloader_params:
shuffle: True
batch_size: 256
num_workers: 8
drop_last: False
pin_memory: True
val_dataset_params:
root: ./data/cifar10
train: False
transforms:
- Resize:
size: 32
- ToTensor
- Normalize:
mean:
- 0.4914
- 0.4822
- 0.4465
std:
- 0.2023
- 0.1994
- 0.2010
target_transform: null
download: True
val_dataloader_params:
batch_size: 512
num_workers: 8
drop_last: False
pin_memory: True
---
Src/Super Gradients/Recipes/Dataset Params/Cityscapes Al Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/cityscapes_al_dataset_params.yaml)
Cityscapes AutoLabelled dataset were introduced by NVIDIA research group.
paper:
Hierarchical Multi-Scale Attention for Semantic Segmentation", https://arxiv.org/abs/2005.10821
Official repo:
https://github.com/NVIDIA/semantic-segmentation
#
AutoLabelled refer to the refinement of the Cityscapes coarse data and pseudo labels generation using their suggested
Hierarchical multi-scale attention model.
#
For dataset preparation instruction please follow:
https://github.com/Deci-AI/super-gradients/blob/master/src/super_gradients/training/datasets/Dataset_Setup_Instructions.md
train_dataset_params:
root_dir: /data/cityscapes
labels_csv_path: lists/labels.csv
list_files:
- lists/train.lst
- lists/auto_labelling.lst
cache_labels: False
cache_images: False
transforms:
- SegStandardize:
max_value: 255
- SegNormalize:
mean: [0.485, 0.456, 0.406]
std: [0.229, 0.224, 0.225]
- SegConvertToTensor:
mask_output_dtype: long
val_dataset_params:
root_dir: /data/cityscapes
list_file: lists/val.lst
labels_csv_path: lists/labels.csv
cache_labels: False
cache_images: False
transforms:
- SegStandardize:
max_value: 255
- SegNormalize:
mean: [0.485, 0.456, 0.406]
std: [0.229, 0.224, 0.225]
- SegConvertToTensor:
mask_output_dtype: long
train_dataloader_params:
dataset: CityscapesConcatDataset
shuffle: True
batch_size: 8
num_workers: 8
drop_last: True # drop the last incomplete batch, if dataset size is not divisible by the batch size
val_dataloader_params:
dataset: CityscapesDataset
batch_size: 8
num_workers: 8
drop_last: False
---
Src/Super Gradients/Recipes/Dataset Params/Cityscapes Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/cityscapes_dataset_params.yaml)
train_dataset_params:
root_dir: /data/cityscapes
list_file: lists/train.lst
labels_csv_path: lists/labels.csv
cache_labels: False
cache_images: False
transforms:
- SegStandardize:
max_value: 255
- SegNormalize:
mean: [0.485, 0.456, 0.406]
std: [0.229, 0.224, 0.225]
- SegConvertToTensor:
mask_output_dtype: long
val_dataset_params:
root_dir: /data/cityscapes
list_file: lists/val.lst
labels_csv_path: lists/labels.csv
cache_labels: False
cache_images: False
transforms:
- SegStandardize:
max_value: 255
- SegNormalize:
mean: [0.485, 0.456, 0.406]
std: [0.229, 0.224, 0.225]
- SegConvertToTensor:
mask_output_dtype: long
train_dataloader_params:
shuffle: True
batch_size: 8
num_workers: 8
drop_last: True # drop the last incomplete batch, if dataset size is not divisible by the batch size
val_dataloader_params:
batch_size: 8
num_workers: 8
drop_last: False
---
Src/Super Gradients/Recipes/Dataset Params/Cityscapes Ddrnet Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/cityscapes_ddrnet_dataset_params.yaml)
defaults:
- cityscapes_dataset_params
- _self_
train_dataset_params:
transforms:
- SegColorJitter:
brightness: 0.5
contrast: 0.5
saturation: 0.5
- SegRandomFlip:
prob: 0.5
- SegRandomRescale:
scales: [ 0.5, 2. ]
- SegPadShortToCropSize:
crop_size: [ 1024, 1024 ]
fill_mask: 19
- SegCropImageAndMask:
crop_size: [ 1024, 1024 ]
mode: random
- SegStandardize:
max_value: 255
- SegNormalize:
mean: [0.485, 0.456, 0.406]
std: [0.229, 0.224, 0.225]
- SegConvertToTensor:
mask_output_dtype: long
train_dataloader_params:
batch_size: 6
val_dataloader_params:
batch_size: 6
---
Src/Super Gradients/Recipes/Dataset Params/Cityscapes Ppliteseg Seg75 Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/cityscapes_ppliteseg_seg75_dataset_params.yaml)
defaults:
- cityscapes_dataset_params
- _self_
train_dataset_params:
transforms:
# for more options see common.factories.transforms_factory.py
- SegColorJitter:
brightness: 0.5
contrast: 0.5
saturation: 0.5
- SegRandomFlip:
prob: 0.5
- SegRandomRescale:
scales: [ 0.25, 1.75 ]
- SegPadShortToCropSize:
crop_size: [ 768, 768 ]
fill_mask: 19 # ignored label idx
- SegCropImageAndMask:
crop_size: [ 768, 768 ]
mode: random
- SegStandardize:
max_value: 255
- SegNormalize:
mean: [0.485, 0.456, 0.406]
std: [0.229, 0.224, 0.225]
- SegConvertToTensor:
mask_output_dtype: long
val_dataset_params:
transforms:
- SegRescale:
scale_factor: 0.75
- SegStandardize:
max_value: 255
- SegNormalize:
mean: [0.485, 0.456, 0.406]
std: [0.229, 0.224, 0.225]
- SegConvertToTensor:
mask_output_dtype: long
train_dataloader_params:
batch_size: 8
val_dataloader_params:
batch_size: 4
---
Src/Super Gradients/Recipes/Dataset Params/Cityscapes Regseg48 Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/cityscapes_regseg48_dataset_params.yaml)
defaults:
- cityscapes_dataset_params
- _self_
train_dataset_params:
transforms:
# for more options see common.factories.transforms_factory.py
- SegColorJitter:
brightness: 0.1
contrast: 0.1
saturation: 0.1
- SegRandomFlip:
prob: 0.5
- SegRandomRescale:
scales: [ 0.4, 1.6 ]
- SegPadShortToCropSize:
crop_size: 1024
fill_image: [ 19, 0, 0 ]
fill_mask: 19 # ignored label idx
- SegCropImageAndMask:
crop_size: 1024
mode: random
- SegStandardize:
max_value: 255
- SegNormalize:
mean: [0.485, 0.456, 0.406]
std: [0.229, 0.224, 0.225]
- SegConvertToTensor:
mask_output_dtype: long
train_dataloader_params:
batch_size: 4
num_workers: 0
val_dataloader_params:
batch_size: 4
num_workers: 0
---
Src/Super Gradients/Recipes/Dataset Params/Cityscapes Segformer Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/cityscapes_segformer_dataset_params.yaml)
defaults:
- cityscapes_dataset_params
- _self_
train_dataset_params:
transforms:
- SegColorJitter:
brightness: 0.5
contrast: 0.5
saturation: 0.5
- SegRandomFlip:
prob: 0.5
- SegRandomRescale:
scales: [ 0.5, 2.0 ]
- SegPadShortToCropSize:
crop_size: [ 1024, 1024 ]
fill_mask: 19
- SegCropImageAndMask:
crop_size: [ 1024, 1024 ]
mode: random
- SegStandardize:
max_value: 255
- SegNormalize:
mean: [0.485, 0.456, 0.406]
std: [0.229, 0.224, 0.225]
- SegConvertToTensor:
mask_output_dtype: long
val_dataset_params:
transforms:
- SegRescale:
long_size: 1024
- SegPadShortToCropSize:
crop_size: [ 1024, 1024 ]
fill_mask: 19
- SegStandardize:
max_value: 255
- SegNormalize:
mean: [0.485, 0.456, 0.406]
std: [0.229, 0.224, 0.225]
- SegConvertToTensor:
mask_output_dtype: long
train_dataloader_params:
batch_size: 2
shuffle: True
val_dataloader_params:
batch_size: 2
shuffle: False
---
Src/Super Gradients/Recipes/Dataset Params/Cityscapes Stdc Seg50 Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/cityscapes_stdc_seg50_dataset_params.yaml)
defaults:
- cityscapes_dataset_params
- _self_
train_dataset_params:
transforms:
# for more options see common.factories.transforms_factory.py
- SegColorJitter:
brightness: 0.5
contrast: 0.5
saturation: 0.5
- SegRandomFlip:
prob: 0.5
- SegRandomRescale:
scales: [ 0.125, 1.5 ]
- SegPadShortToCropSize:
crop_size: [ 1024, 512 ]
fill_mask: 19 # ignored label idx
- SegCropImageAndMask:
crop_size: [ 1024, 512 ]
mode: random
- SegStandardize:
max_value: 255
- SegNormalize:
mean: [0.485, 0.456, 0.406]
std: [0.229, 0.224, 0.225]
- SegConvertToTensor:
mask_output_dtype: long
val_dataset_params:
transforms:
- SegRescale:
scale_factor: 0.5
- SegStandardize:
max_value: 255
- SegNormalize:
mean: [0.485, 0.456, 0.406]
std: [0.229, 0.224, 0.225]
- SegConvertToTensor:
mask_output_dtype: long
train_dataloader_params:
batch_size: 16
val_dataloader_params:
batch_size: 16
---
Src/Super Gradients/Recipes/Dataset Params/Cityscapes Stdc Seg75 Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/cityscapes_stdc_seg75_dataset_params.yaml)
defaults:
- cityscapes_dataset_params
- _self_
train_dataset_params:
transforms:
# for more options see common.factories.transforms_factory.py
- SegColorJitter:
brightness: 0.5
contrast: 0.5
saturation: 0.5
- SegRandomFlip:
prob: 0.5
- SegRandomRescale:
scales: [ 0.125, 1.5 ]
- SegPadShortToCropSize:
crop_size: [ 1536, 768 ]
fill_mask: 19 # ignored label idx
- SegCropImageAndMask:
crop_size: [ 1536, 768 ]
mode: random
- SegStandardize:
max_value: 255
- SegNormalize:
mean: [0.485, 0.456, 0.406]
std: [0.229, 0.224, 0.225]
- SegConvertToTensor:
mask_output_dtype: long
val_dataset_params:
transforms:
- SegRescale:
scale_factor: 0.75
- SegStandardize:
max_value: 255
- SegNormalize:
mean: [ 0.485, 0.456, 0.406 ]
std: [ 0.229, 0.224, 0.225 ]
- SegConvertToTensor:
mask_output_dtype: long
train_dataloader_params:
batch_size: 4
val_dataloader_params:
batch_size: 4
---
Src/Super Gradients/Recipes/Dataset Params/Coco Detection Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/coco_detection_dataset_params.yaml)
train_dataset_params:
data_dir: /data/coco # root path to coco data
subdir: images/train2017 # sub directory path of data_dir containing the train data.
json_file: instances_train2017.json # path to coco train json file, data_dir/annotations/train_json_file.
input_dim: [640, 640]
cache_annotations: True
ignore_empty_annotations: True
transforms:
- DetectionMosaic:
input_dim: ${dataset_params.train_dataset_params.input_dim}
prob: 1.
- DetectionRandomAffine:
degrees: 10. # rotation degrees, randomly sampled from [-degrees, degrees]
translate: 0.1 # image translation fraction
scales: [ 0.1, 2 ] # random rescale range (keeps size by padding/cropping) after mosaic transform.
shear: 2.0 # shear degrees, randomly sampled from [-degrees, degrees]
target_size: ${dataset_params.train_dataset_params.input_dim}
filter_box_candidates: True # whether to filter out transformed bboxes by edge size, area ratio, and aspect ratio.
wh_thr: 2 # edge size threshold when filter_box_candidates = True (pixels)
area_thr: 0.1 # threshold for area ratio between original image and the transformed one, when when filter_box_candidates = True
ar_thr: 20 # aspect ratio threshold when filter_box_candidates = True
- DetectionMixup:
input_dim: ${dataset_params.train_dataset_params.input_dim}
mixup_scale: [ 0.5, 1.5 ] # random rescale range for the additional sample in mixup
prob: 1.0 # probability to apply per-sample mixup
flip_prob: 0.5 # probability to apply horizontal flip
- DetectionHSV:
prob: 1.0 # probability to apply HSV transform
hgain: 5 # HSV transform hue gain (randomly sampled from [-hgain, hgain])
sgain: 30 # HSV transform saturation gain (randomly sampled from [-sgain, sgain])
vgain: 30 # HSV transform value gain (randomly sampled from [-vgain, vgain])
- DetectionHorizontalFlip:
prob: 0.5 # probability to apply horizontal flip
- DetectionPaddedRescale:
input_dim: ${dataset_params.train_dataset_params.input_dim}
- DetectionTargetsFormatTransform:
input_dim: ${dataset_params.train_dataset_params.input_dim}
output_format: LABEL_CXCYWH
class_inclusion_list:
max_num_samples:
with_crowd: False
train_dataloader_params:
shuffle: True
batch_size: 16
num_workers: 8
drop_last: True
pin_memory: True
worker_init_fn:
_target_: super_gradients.training.utils.utils.load_func
dotpath: super_gradients.training.datasets.datasets_utils.worker_init_reset_seed
collate_fn: DetectionCollateFN
val_dataset_params:
data_dir: /data/coco # root path to coco data
subdir: images/val2017 # sub directory path of data_dir containing the train data.
json_file: instances_val2017.json # path to coco train json file, data_dir/annotations/train_json_file.
input_dim: [640, 640]
cache_annotations: True
ignore_empty_annotations: True
transforms:
- DetectionPaddedRescale:
input_dim: ${dataset_params.val_dataset_params.input_dim}
- DetectionTargetsFormatTransform:
input_dim: ${dataset_params.val_dataset_params.input_dim}
output_format: LABEL_CXCYWH
class_inclusion_list:
max_num_samples:
with_crowd: True
val_dataloader_params:
batch_size: 64
num_workers: 8
drop_last: False
pin_memory: True
collate_fn: CrowdDetectionCollateFN
_convert_: all
---
Src/Super Gradients/Recipes/Dataset Params/Coco Detection Ppyoloe Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/coco_detection_ppyoloe_dataset_params.yaml)
train_dataset_params:
data_dir: /data/coco # root path to coco data
subdir: images/train2017 # sub directory path of data_dir containing the train data.
json_file: instances_train2017.json # path to coco train json file, data_dir/annotations/train_json_file.
input_dim: # None, do not resize dataset on load
cache_annotations: True
ignore_empty_annotations: True
transforms:
- DetectionRandomAffine:
degrees: 0 # rotation degrees, randomly sampled from [-degrees, degrees]
translate: 0.25 # image translation fraction
scales: [ 0.5, 1.5 ] # random rescale range (keeps size by padding/cropping) after mosaic transform.
shear: 0.0 # shear degrees, randomly sampled from [-degrees, degrees]
target_size:
filter_box_candidates: True # whether to filter out transformed bboxes by edge size, area ratio, and aspect ratio.
wh_thr: 2 # edge size threshold when filter_box_candidates = True (pixels)
area_thr: 0.1 # threshold for area ratio between original image and the transformed one, when when filter_box_candidates = True
ar_thr: 20 # aspect ratio threshold when filter_box_candidates = True
- DetectionRandomRotate90:
prob: 0.5
- DetectionRGB2BGR:
prob: 0.25
- DetectionHSV:
prob: 0.5 # probability to apply HSV transform
hgain: 18 # HSV transform hue gain (randomly sampled from [-hgain, hgain])
sgain: 30 # HSV transform saturation gain (randomly sampled from [-sgain, sgain])
vgain: 30 # HSV transform value gain (randomly sampled from [-vgain, vgain])
- DetectionHorizontalFlip:
prob: 0.5 # probability to apply horizontal flip
- DetectionMixup:
input_dim:
mixup_scale: [ 0.5, 1.5 ] # random rescale range for the additional sample in mixup
prob: 0.5 # probability to apply per-sample mixup
flip_prob: 0.5 # probability to apply horizontal flip
- DetectionNormalize:
mean: [ 123.675, 116.28, 103.53 ]
std: [ 58.395, 57.12, 57.375 ]
- DetectionTargetsFormatTransform:
output_format: LABEL_CXCYWH
class_inclusion_list:
max_num_samples:
with_crowd: False
train_dataloader_params:
batch_size: 32
num_workers: 8
shuffle: True
drop_last: True
# Disable pin_memory due to presence of PPYoloECollateFN with uses random resize during training
pin_memory: False
worker_init_fn:
_target_: super_gradients.training.utils.utils.load_func
dotpath: super_gradients.training.datasets.datasets_utils.worker_init_reset_seed
collate_fn:
PPYoloECollateFN:
random_resize_sizes: [ 320, 352, 384, 416, 448, 480, 512, 544, 576, 608, 640, 672, 704, 736, 768 ]
random_resize_modes:
- 0 # cv::INTER_NEAREST
- 1 # cv::INTER_LINEAR
- 2 # cv::INTER_CUBIC
- 3 # cv::INTER_AREA
- 4 # cv::INTER_LANCZOS4
val_dataset_params:
data_dir: /data/coco # root path to coco data
subdir: images/val2017 # sub directory path of data_dir containing the train data.
json_file: instances_val2017.json # path to coco train json file, data_dir/annotations/train_json_file.
input_dim:
cache_annotations: True
ignore_empty_annotations: True
transforms:
- DetectionRescale:
output_shape: [640, 640]
- DetectionNormalize:
mean: [ 123.675, 116.28, 103.53 ]
std: [ 58.395, 57.12, 57.375 ]
- DetectionTargetsFormatTransform:
output_format: LABEL_CXCYWH
class_inclusion_list:
max_num_samples:
with_crowd: True
val_dataloader_params:
batch_size: 64
num_workers: 8
drop_last: False
shuffle: False
pin_memory: False
collate_fn: CrowdDetectionPPYoloECollateFN
_convert_: all
---
Src/Super Gradients/Recipes/Dataset Params/Coco Detection Ssd Lite Mobilenet V2 Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/coco_detection_ssd_lite_mobilenet_v2_dataset_params.yaml)
defaults:
- coco_detection_dataset_params
train_dataset_params:
data_dir: /data/coco # root path to coco data
subdir: images/train2017 # sub directory path of data_dir containing the train data.
json_file: instances_train2017.json # path to coco train json file, data_dir/annotations/train_json_file.
input_dim: [320, 320]
cache_annotations: True
ignore_empty_annotations: True
transforms:
- DetectionRandomAffine:
degrees: 0. # rotation degrees, randomly sampled from [-degrees, degrees]
translate: 0.1 # image translation fraction
scales: [0.5, 1.5] # random rescale range (keeps size by padding/cropping) after mosaic transform.
shear: 0. # shear degrees, randomly sampled from [-degrees, degrees]
target_size: ${dataset_params.train_dataset_params.input_dim}
filter_box_candidates: True # whether to filter out transformed bboxes by edge size, area ratio, and aspect ratio.
wh_thr: 2 # edge size threshold when filter_box_candidates = True (pixels)
area_thr: 0.1 # threshold for area ratio between original image and the transformed one, when when filter_box_candidates = True
ar_thr: 20 # aspect ratio threshold when filter_box_candidates = True
- DetectionHSV:
prob: 1.0 # probability to apply HSV transform
hgain: 5 # HSV transform hue gain (randomly sampled from [-hgain, hgain])
sgain: 30 # HSV transform saturation gain (randomly sampled from [-sgain, sgain])
vgain: 30 # HSV transform value gain (randomly sampled from [-vgain, vgain])
- DetectionHorizontalFlip:
prob: 0.5 # probability to apply horizontal flip
- DetectionPaddedRescale:
input_dim: ${dataset_params.train_dataset_params.input_dim}
- DetectionTargetsFormatTransform:
input_dim: ${dataset_params.train_dataset_params.input_dim}
output_format: LABEL_NORMALIZED_CXCYWH
class_inclusion_list:
max_num_samples:
with_crowd: False
train_dataloader_params:
batch_size: 32
num_workers: 8
shuffle: True
drop_last: True
pin_memory: True
worker_init_fn:
_target_: super_gradients.training.utils.utils.load_func
dotpath: super_gradients.training.datasets.datasets_utils.worker_init_reset_seed
collate_fn: DetectionCollateFN
val_dataset_params:
data_dir: /data/coco # root path to coco data
subdir: images/val2017 # sub directory path of data_dir containing the train data.
json_file: instances_val2017.json # path to coco train json file, data_dir/annotations/train_json_file.
input_dim: [320, 320]
cache_annotations: True
ignore_empty_annotations: True
transforms:
- DetectionPaddedRescale:
input_dim: ${dataset_params.val_dataset_params.input_dim}
- DetectionTargetsFormatTransform:
input_dim: ${dataset_params.val_dataset_params.input_dim}
output_format: LABEL_NORMALIZED_CXCYWH
class_inclusion_list:
max_num_samples:
with_crowd: True
val_dataloader_params:
batch_size: 16
num_workers: 8
drop_last: False
pin_memory: True
collate_fn: CrowdDetectionCollateFN
_convert_: all
---
Src/Super Gradients/Recipes/Dataset Params/Coco Detection Yolo Format Base Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/coco_detection_yolo_format_base_dataset_params.yaml)
train_dataset_params:
data_dir: /data/coco # TO FILL: Where the data is stored.
images_dir: images/train2017 # TO FILL: Local path to directory that includes all the images. Path relative to data_dir. Can be the same as labels_dir.
labels_dir: labels/train2017 # TO FILL: Local path to directory that includes all the labels. Path relative to data_dir. Can be the same as images_dir.
classes: [ person, bicycle, car, motorcycle, airplane, bus, train, truck, boat, traffic light, fire hydrant, stop sign,
parking meter, bench, bird, cat, dog, horse, sheep, cow, elephant, bear, zebra, giraffe, backpack, umbrella, handbag,
tie, suitcase, frisbee, skis, snowboard, sports ball, kite, baseball bat, baseball glove, skateboard, surfboard,
tennis racket, bottle, wine glass, cup, fork, knife, spoon, bowl, banana, apple, sandwich, orange, broccoli, carrot,
hot dog, pizza, donut, cake, chair, couch, potted plant, bed, dining table, toilet, tv, laptop, mouse, remote,
keyboard, cell phone, microwave, oven, toaster, sink, refrigerator, book, clock, vase, scissors, teddy bear,
hair drier, toothbrush] # TO FILL: List of classes used in your dataset.
input_dim: [640, 640]
cache_annotations: True
ignore_empty_annotations: True
transforms:
- DetectionMosaic:
input_dim: ${dataset_params.train_dataset_params.input_dim}
prob: 1.
- DetectionRandomAffine:
degrees: 10. # rotation degrees, randomly sampled from [-degrees, degrees]
translate: 0.1 # image translation fraction
scales: [ 0.1, 2 ] # random rescale range (keeps size by padding/cropping) after mosaic transform.
shear: 2.0 # shear degrees, randomly sampled from [-degrees, degrees]
target_size: ${dataset_params.train_dataset_params.input_dim}
filter_box_candidates: True # whether to filter out transformed bboxes by edge size, area ratio, and aspect ratio.
wh_thr: 2 # edge size threshold when filter_box_candidates = True (pixels)
area_thr: 0.1 # threshold for area ratio between original image and the transformed one, when when filter_box_candidates = True
ar_thr: 20 # aspect ratio threshold when filter_box_candidates = True
- DetectionMixup:
input_dim: ${dataset_params.train_dataset_params.input_dim}
mixup_scale: [ 0.5, 1.5 ] # random rescale range for the additional sample in mixup
prob: 1.0 # probability to apply per-sample mixup
flip_prob: 0.5 # probability to apply horizontal flip
- DetectionHSV:
prob: 1.0 # probability to apply HSV transform
hgain: 5 # HSV transform hue gain (randomly sampled from [-hgain, hgain])
sgain: 30 # HSV transform saturation gain (randomly sampled from [-sgain, sgain])
vgain: 30 # HSV transform value gain (randomly sampled from [-vgain, vgain])
- DetectionHorizontalFlip:
prob: 0.5 # probability to apply horizontal flip
- DetectionPaddedRescale:
input_dim: ${dataset_params.train_dataset_params.input_dim}
- DetectionTargetsFormatTransform:
input_dim: ${dataset_params.train_dataset_params.input_dim}
output_format: LABEL_CXCYWH
class_inclusion_list:
max_num_samples:
train_dataloader_params:
batch_size: 25
num_workers: 8
shuffle: True
drop_last: True
pin_memory: True
collate_fn: DetectionCollateFN
val_dataset_params:
data_dir: /data/coco # TO FILL: Where the data is stored.
images_dir: images/val2017 # TO FILL: Local path to directory that includes all the images. Path relative to data_dir. Can be the same as labels_dir.
labels_dir: labels/val2017 # TO FILL: Local path to directory that includes all the labels. Path relative to data_dir. Can be the same as images_dir.
classes: [ person, bicycle, car, motorcycle, airplane, bus, train, truck, boat, traffic light, fire hydrant, stop sign,
parking meter, bench, bird, cat, dog, horse, sheep, cow, elephant, bear, zebra, giraffe, backpack, umbrella, handbag,
tie, suitcase, frisbee, skis, snowboard, sports ball, kite, baseball bat, baseball glove, skateboard, surfboard,
tennis racket, bottle, wine glass, cup, fork, knife, spoon, bowl, banana, apple, sandwich, orange, broccoli, carrot,
hot dog, pizza, donut, cake, chair, couch, potted plant, bed, dining table, toilet, tv, laptop, mouse, remote,
keyboard, cell phone, microwave, oven, toaster, sink, refrigerator, book, clock, vase, scissors, teddy bear,
hair drier, toothbrush] # TO FILL: List of classes used in your dataset.
input_dim: [640, 640]
cache_annotations: True
ignore_empty_annotations: True
transforms:
- DetectionPaddedRescale:
input_dim: ${dataset_params.val_dataset_params.input_dim}
- DetectionTargetsFormatTransform:
input_dim: ${dataset_params.val_dataset_params.input_dim}
output_format: LABEL_CXCYWH
class_inclusion_list:
max_num_samples:
val_dataloader_params:
batch_size: 25
num_workers: 8
drop_last: False
pin_memory: True
collate_fn: DetectionCollateFN
_convert_: all
---
Src/Super Gradients/Recipes/Dataset Params/Coco Detection Yolo Nas Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/coco_detection_yolo_nas_dataset_params.yaml)
class_names: [
"person",
"bicycle",
"car",
"motorcycle",
"airplane",
"bus",
"train",
"truck",
"boat",
"traffic light",
"fire hydrant",
"stop sign",
"parking meter",
"bench",
"bird",
"cat",
"dog",
"horse",
"sheep",
"cow",
"elephant",
"bear",
"zebra",
"giraffe",
"backpack",
"umbrella",
"handbag",
"tie",
"suitcase",
"frisbee",
"skis",
"snowboard",
"sports ball",
"kite",
"baseball bat",
"baseball glove",
"skateboard",
"surfboard",
"tennis racket",
"bottle",
"wine glass",
"cup",
"fork",
"knife",
"spoon",
"bowl",
"banana",
"apple",
"sandwich",
"orange",
"broccoli",
"carrot",
"hot dog",
"pizza",
"donut",
"cake",
"chair",
"couch",
"potted plant",
"bed",
"dining table",
"toilet",
"tv",
"laptop",
"mouse",
"remote",
"keyboard",
"cell phone",
"microwave",
"oven",
"toaster",
"sink",
"refrigerator",
"book",
"clock",
"vase",
"scissors",
"teddy bear",
"hair drier",
"toothbrush",
]
train_dataset_params:
data_dir: /data/coco # root path to coco data
subdir: images/train2017 # sub directory path of data_dir containing the train data.
json_file: instances_train2017.json # path to coco train json file, data_dir/annotations/train_json_file.
input_dim: [640, 640]
cache_annotations: True
ignore_empty_annotations: True
transforms:
- DetectionRandomAffine:
degrees: 0 # rotation degrees, randomly sampled from [-degrees, degrees]
translate: 0.25 # image translation fraction
scales: [ 0.5, 1.5 ] # random rescale range (keeps size by padding/cropping) after mosaic transform.
shear: 0.0 # shear degrees, randomly sampled from [-degrees, degrees]
target_size:
filter_box_candidates: True # whether to filter out transformed bboxes by edge size, area ratio, and aspect ratio.
wh_thr: 2 # edge size threshold when filter_box_candidates = True (pixels)
area_thr: 0.1 # threshold for area ratio between original image and the transformed one, when when filter_box_candidates = True
ar_thr: 20 # aspect ratio threshold when filter_box_candidates = True
- DetectionRGB2BGR:
prob: 0.5
- DetectionHSV:
prob: 0.5 # probability to apply HSV transform
hgain: 18 # HSV transform hue gain (randomly sampled from [-hgain, hgain])
sgain: 30 # HSV transform saturation gain (randomly sampled from [-sgain, sgain])
vgain: 30 # HSV transform value gain (randomly sampled from [-vgain, vgain])
- DetectionHorizontalFlip:
prob: 0.5 # probability to apply horizontal flip
- DetectionMixup:
input_dim:
mixup_scale: [ 0.5, 1.5 ] # random rescale range for the additional sample in mixup
prob: 0.5 # probability to apply per-sample mixup
flip_prob: 0.5 # probability to apply horizontal flip
- DetectionPaddedRescale:
input_dim: ${dataset_params.train_dataset_params.input_dim}
pad_value: 114
- DetectionStandardize:
max_value: 255.
- DetectionTargetsFormatTransform:
output_format: LABEL_CXCYWH
class_inclusion_list:
max_num_samples:
with_crowd: False
train_dataloader_params:
batch_size: 25
num_workers: 8
shuffle: True
drop_last: True
pin_memory: True
collate_fn: DetectionCollateFN
val_dataset_params:
data_dir: /data/coco # root path to coco data
subdir: images/val2017 # sub directory path of data_dir containing the train data.
json_file: instances_val2017.json # path to coco train json file, data_dir/annotations/train_json_file.
input_dim: [636, 636]
cache_annotations: True
ignore_empty_annotations: True
transforms:
- DetectionRGB2BGR:
prob: 1
- DetectionPadToSize:
output_size: [640, 640]
pad_value: 114
- DetectionStandardize:
max_value: 255.
- DetectionImagePermute
- DetectionTargetsFormatTransform:
input_dim: [640, 640]
output_format: LABEL_CXCYWH
class_inclusion_list:
max_num_samples:
with_crowd: True
val_dataloader_params:
batch_size: 25
num_workers: 8
drop_last: False
shuffle: False
pin_memory: True
collate_fn: CrowdDetectionCollateFN
_convert_: all
---
Src/Super Gradients/Recipes/Dataset Params/Coco Pose Common Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/coco_pose_common_dataset_params.yaml)
This file is not "true" dataset_params file, but rather a collection of settings that describe
skeleton configuration specific to COCO dataset. It is used by other dataset_params files to
avoid code duplication.
num_joints: 17
OKs sigma values take from https://github.com/cocodataset/cocoapi/blob/master/PythonAPI/pycocotools/cocoeval.py#L523
oks_sigmas: [0.026, 0.025, 0.025, 0.035, 0.035, 0.079, 0.079, 0.072, 0.072, 0.062, 0.062, 0.107, 0.107, 0.087, 0.087, 0.089, 0.089]
flip_indexes: [ 0, 2, 1, 4, 3, 6, 5, 8, 7, 10, 9, 12, 11, 14, 13, 16, 15,]
edge_links:
- [0, 1]
- [0, 2]
- [1, 2]
- [1, 3]
- [2, 4]
- [3, 5]
- [4, 6]
- [5, 6]
- [5, 7]
- [5, 11]
- [6, 8]
- [6, 12]
- [7, 9]
- [8, 10]
- [11, 12]
- [11, 13]
- [12, 14]
- [13, 15]
- [14, 16]
edge_colors:
- [214, 39, 40] # Nose -> LeftEye
- [148, 103, 189] # Nose -> RightEye
- [44, 160, 44] # LeftEye -> RightEye
- [140, 86, 75] # LeftEye -> LeftEar
- [227, 119, 194] # RightEye -> RightEar
- [127, 127, 127] # LeftEar -> LeftShoulder
- [188, 189, 34] # RightEar -> RightShoulder
- [127, 127, 127] # Shoulders
- [188, 189, 34] # LeftShoulder -> LeftElbow
- [140, 86, 75] # LeftTorso
- [23, 190, 207] # RightShoulder -> RightElbow
- [227, 119, 194] # RightTorso
- [31, 119, 180] # LeftElbow -> LeftArm
- [255, 127, 14] # RightElbow -> RightArm
- [148, 103, 189] # Waist
- [255, 127, 14] # Left Hip -> Left Knee
- [214, 39, 40] # Right Hip -> Right Knee
- [31, 119, 180] # Left Knee -> Left Ankle
- [44, 160, 44] # Right Knee -> Right Ankle
keypoint_colors:
- [148, 103, 189]
- [31, 119, 180]
- [148, 103, 189]
- [31, 119, 180]
- [148, 103, 189]
- [31, 119, 180]
- [148, 103, 189]
- [31, 119, 180]
- [148, 103, 189]
- [31, 119, 180]
- [148, 103, 189]
- [31, 119, 180]
- [148, 103, 189]
- [31, 119, 180]
- [148, 103, 189]
- [31, 119, 180]
- [148, 103, 189]
---
Src/Super Gradients/Recipes/Dataset Params/Coco Pose Estimation Common Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/coco_pose_estimation_common_dataset_params.yaml)
This is not "true" dataset params, one cannot use it to instantiate dataloaders
But it contains skeleton definitions for COCO2017 dataset and exists to avoid
duplication of those parameters in other dataset params
num_joints: 17
OKs sigma values take from https://github.com/cocodataset/cocoapi/blob/master/PythonAPI/pycocotools/cocoeval.py#L523
oks_sigmas: [0.026, 0.025, 0.025, 0.035, 0.035, 0.079, 0.079, 0.072, 0.072, 0.062, 0.062, 0.107, 0.107, 0.087, 0.087, 0.089, 0.089]
flip_indexes: [ 0, 2, 1, 4, 3, 6, 5, 8, 7, 10, 9, 12, 11, 14, 13, 16, 15,]
edge_links:
- [0, 1]
- [0, 2]
- [1, 2]
- [1, 3]
- [2, 4]
- [3, 5]
- [4, 6]
- [5, 6]
- [5, 7]
- [5, 11]
- [6, 8]
- [6, 12]
- [7, 9]
- [8, 10]
- [11, 12]
- [11, 13]
- [12, 14]
- [13, 15]
- [14, 16]
edge_colors:
- [214, 39, 40] # Nose -> LeftEye
- [148, 103, 189] # Nose -> RightEye
- [44, 160, 44] # LeftEye -> RightEye
- [140, 86, 75] # LeftEye -> LeftEar
- [227, 119, 194] # RightEye -> RightEar
- [127, 127, 127] # LeftEar -> LeftShoulder
- [188, 189, 34] # RightEar -> RightShoulder
- [127, 127, 127] # Shoulders
- [188, 189, 34] # LeftShoulder -> LeftElbow
- [140, 86, 75] # LeftTorso
- [23, 190, 207] # RightShoulder -> RightElbow
- [227, 119, 194] # RightTorso
- [31, 119, 180] # LeftElbow -> LeftArm
- [255, 127, 14] # RightElbow -> RightArm
- [148, 103, 189] # Waist
- [255, 127, 14] # Left Hip -> Left Knee
- [214, 39, 40] # Right Hip -> Right Knee
- [31, 119, 180] # Left Knee -> Left Ankle
- [44, 160, 44] # Right Knee -> Right Ankle
keypoint_colors:
- [148, 103, 189]
- [31, 119, 180]
- [148, 103, 189]
- [31, 119, 180]
- [148, 103, 189]
- [31, 119, 180]
- [148, 103, 189]
- [31, 119, 180]
- [148, 103, 189]
- [31, 119, 180]
- [148, 103, 189]
- [31, 119, 180]
- [148, 103, 189]
- [31, 119, 180]
- [148, 103, 189]
- [31, 119, 180]
- [148, 103, 189]
---
Src/Super Gradients/Recipes/Dataset Params/Coco Pose Estimation Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/coco_pose_estimation_dataset_params.yaml)
defaults:
- coco_pose_estimation_common_dataset_params
- _self_
train_dataset_params:
data_dir: /data/coco # root path to coco data
images_dir: images/train2017
json_file: annotations/person_keypoints_train2017.json
include_empty_samples: False
min_instance_area: 64
edge_links: ${dataset_params.edge_links}
edge_colors: ${dataset_params.edge_colors}
keypoint_colors: ${dataset_params.keypoint_colors}
transforms:
- KeypointsLongestMaxSize:
max_height: 640
max_width: 640
- KeypointsPadIfNeeded:
min_height: 640
min_width: 640
image_pad_value: 127
mask_pad_value: 1
- KeypointsRandomHorizontalFlip:
# Note these indexes are COCO-specific. If you're using a different dataset, you'll need to change these accordingly.
flip_index: ${dataset_params.flip_indexes}
prob: 0.5
- KeypointsRandomAffineTransform:
max_rotation: 30
min_scale: 0.5
max_scale: 2
max_translate: 0.2
image_pad_value: 127
mask_pad_value: 1
prob: 0.75
- KeypointsImageStandardize:
max_value: 255
- KeypointsImageNormalize:
mean: [ 0.485, 0.456, 0.406 ]
std: [ 0.229, 0.224, 0.225 ]
- KeypointsImageToTensor
val_dataset_params:
data_dir: /data/coco/
images_dir: images/val2017
json_file: annotations/person_keypoints_val2017.json
include_empty_samples: True
min_instance_area: 128
edge_links: ${dataset_params.edge_links}
edge_colors: ${dataset_params.edge_colors}
keypoint_colors: ${dataset_params.keypoint_colors}
transforms:
- KeypointsLongestMaxSize:
max_height: 640
max_width: 640
- KeypointsPadIfNeeded:
min_height: 640
min_width: 640
image_pad_value: 127
mask_pad_value: 1
- KeypointsImageStandardize:
max_value: 255
- KeypointsImageNormalize:
mean: [ 0.485, 0.456, 0.406 ]
std: [ 0.229, 0.224, 0.225 ]
- KeypointsImageToTensor
train_dataloader_params:
shuffle: True
batch_size: 8
num_workers: 8
drop_last: True
collate_fn: KeypointsCollate
val_dataloader_params:
batch_size: 24
num_workers: 8
drop_last: False
collate_fn: KeypointsCollate
_convert_: all
---
Src/Super Gradients/Recipes/Dataset Params/Coco Pose Estimation Dekr Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/coco_pose_estimation_dekr_dataset_params.yaml)
defaults:
- coco_pose_estimation_dataset_params
- _self_
train_dataset_params:
target_generator:
DEKRTargetsGenerator:
output_stride: 4
sigma: 2
center_sigma: 4
bg_weight: 0.1
offset_radius: 4
val_dataset_params:
target_generator:
DEKRTargetsGenerator:
output_stride: 4
sigma: 2
center_sigma: 4
bg_weight: 0.1
offset_radius: 4
---
Src/Super Gradients/Recipes/Dataset Params/Coco Pose Estimation Rescoring Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/coco_pose_estimation_rescoring_dataset_params.yaml)
defaults:
- coco_pose_estimation_common_dataset_params
- _self_
train_dataset_params:
pkl_file: REPLACE_ME_WITH_PATH_TO_TRAIN_DATASET.PKL
val_dataset_params:
pkl_file: REPLACE_ME_WITH_PATH_TO_VAL_DATASET.PKL
train_dataloader_params:
shuffle: True
batch_size: 1024
num_workers: 0
drop_last: True
val_dataloader_params:
batch_size: 1
num_workers: 0
drop_last: False
_convert_: all
---
Src/Super Gradients/Recipes/Dataset Params/Coco Pose Estimation Yolo Nas Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/coco_pose_estimation_yolo_nas_dataset_params.yaml)
defaults:
- coco_pose_common_dataset_params
- _self_
This is a shortcut parameter to set size of training & validation images.
image_size: 640
dataset_params_suffix: "${dataset_params.image_size}"
train_dataset_params:
data_dir: /data/coco # root path to coco data
images_dir: images/train2017
json_file: annotations/person_keypoints_train2017.json
edge_links: ${dataset_params.edge_links}
edge_colors: ${dataset_params.edge_colors}
keypoint_colors: ${dataset_params.keypoint_colors}
include_empty_samples: True
crowd_annotations_action: mask_as_normal
transforms:
- KeypointsRandomHorizontalFlip:
flip_index: ${dataset_params.flip_indexes}
prob: 0.5
- KeypointsBrightnessContrast:
brightness_range: [ 0.8, 1.2 ]
contrast_range: [ 0.8, 1.2 ]
prob: 0.5
- KeypointsHSV:
hgain: 20
sgain: 20
vgain: 20
prob: 0.5
- KeypointsRandomAffineTransform:
max_rotation: 5
min_scale: 0.5
max_scale: 1.5
max_translate: 0.1
image_pad_value: 127
mask_pad_value: 1
prob: 0.75
interpolation_mode: [0, 1, 2, 3, 4]
- KeypointsLongestMaxSize:
max_height: ${dataset_params.image_size}
max_width: ${dataset_params.image_size}
- KeypointsPadIfNeeded:
min_height: ${dataset_params.image_size}
min_width: ${dataset_params.image_size}
image_pad_value: [127, 127, 127]
mask_pad_value: 1
padding_mode: center
- KeypointsImageStandardize:
max_value: 255
- KeypointsRemoveSmallObjects:
min_instance_area: 1
min_visible_keypoints: 1
val_dataset_params:
data_dir: /data/coco/
images_dir: images/val2017
json_file: annotations/person_keypoints_val2017.json
edge_links: ${dataset_params.edge_links}
edge_colors: ${dataset_params.edge_colors}
keypoint_colors: ${dataset_params.keypoint_colors}
include_empty_samples: True
crowd_annotations_action: no_action
transforms:
- KeypointsLongestMaxSize:
max_height: ${dataset_params.image_size}
max_width: ${dataset_params.image_size}
- KeypointsPadIfNeeded:
min_height: ${dataset_params.image_size}
min_width: ${dataset_params.image_size}
image_pad_value: 127
mask_pad_value: 1
padding_mode: bottom_right
- KeypointsImageStandardize:
max_value: 255
train_dataloader_params:
dataset: COCOPoseEstimationDataset
shuffle: True
batch_size: 8
num_workers: 8
drop_last: True
pin_memory: False
collate_fn: YoloNASPoseCollateFN
val_dataloader_params:
dataset: COCOPoseEstimationDataset
batch_size: 8
num_workers: 8
drop_last: False
pin_memory: False
collate_fn: YoloNASPoseCollateFN
---
Src/Super Gradients/Recipes/Dataset Params/Coco Pose Estimation Yolo Nas Mosaic Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/coco_pose_estimation_yolo_nas_mosaic_dataset_params.yaml)
defaults:
- coco_pose_estimation_yolo_nas_dataset_params
- _self_
mosaic_prob: 0.5
dataset_params_suffix: "mosaic_${dataset_params.mosaic_prob}_${dataset_params.image_size}"
train_dataset_params:
transforms:
- KeypointsRandomHorizontalFlip:
flip_index: ${dataset_params.flip_indexes}
prob: 0.5
- KeypointsBrightnessContrast:
brightness_range: [ 0.8, 1.2 ]
contrast_range: [ 0.8, 1.2 ]
prob: 0.5
- KeypointsHSV:
hgain: 20
sgain: 20
vgain: 20
prob: 0.5
- KeypointsRandomAffineTransform:
max_rotation: 5
min_scale: 0.75
max_scale: 1.5
max_translate: 0.1
image_pad_value: 127
mask_pad_value: 1
prob: 0.75
interpolation_mode: [ 0, 1, 2, 3, 4 ]
- KeypointsMosaic:
prob: ${dataset_params.mosaic_prob}
- KeypointsLongestMaxSize:
max_height: ${dataset_params.image_size}
max_width: ${dataset_params.image_size}
- KeypointsPadIfNeeded:
min_height: ${dataset_params.image_size}
min_width: ${dataset_params.image_size}
image_pad_value: [ 127, 127, 127 ]
mask_pad_value: 1
padding_mode: center
- KeypointsImageStandardize:
max_value: 255
- KeypointsRemoveSmallObjects:
min_instance_area: 1
min_visible_keypoints: 1
---
Src/Super Gradients/Recipes/Dataset Params/Coco Pose Estimation Yolo Nas Mosaic Heavy Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/coco_pose_estimation_yolo_nas_mosaic_heavy_dataset_params.yaml)
defaults:
- coco_pose_estimation_yolo_nas_mosaic_dataset_params
- _self_
dataset_params_suffix: "mosaic_heavy_augs_${dataset_params.mosaic_prob}_${dataset_params.image_size}"
train_dataset_params:
transforms:
- KeypointsRandomHorizontalFlip:
flip_index: ${dataset_params.flip_indexes}
prob: 0.5
- KeypointsBrightnessContrast:
brightness_range: [ 0.7, 1.3 ]
contrast_range: [ 0.7, 1.3 ]
prob: 0.75
- KeypointsReverseImageChannels:
prob: 0.5
- KeypointsHSV:
hgain: 25
sgain: 25
vgain: 25
prob: 0.75
- KeypointsRandomRotate90:
prob: 0.5
- KeypointsRandomAffineTransform:
max_rotation: 7
min_scale: 0.6
max_scale: 1.75
max_translate: 0.1
image_pad_value: 127
mask_pad_value: 1
prob: 0.75
interpolation_mode: [ 0, 1, 2, 3, 4 ]
- KeypointsMosaic:
prob: ${dataset_params.mosaic_prob}
- KeypointsLongestMaxSize:
max_height: ${dataset_params.image_size}
max_width: ${dataset_params.image_size}
- KeypointsPadIfNeeded:
min_height: ${dataset_params.image_size}
min_width: ${dataset_params.image_size}
image_pad_value: [ 127, 127, 127 ]
mask_pad_value: 1
padding_mode: center
- KeypointsImageStandardize:
max_value: 255
- KeypointsRemoveSmallObjects:
min_instance_area: 1
min_visible_keypoints: 1
---
Src/Super Gradients/Recipes/Dataset Params/Coco Segmentation Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/coco_segmentation_dataset_params.yaml)
train_dataset_params:
root_dir: /data/coco/
list_file: instances_train2017.json
samples_sub_directory: images/train2017
targets_sub_directory: annotations
dataset_classes_inclusion_tuples_list:
_target_: super_gradients.training.utils.segmentation_utils.coco_sub_classes_inclusion_tuples_list
cache_labels: False
cache_images: False
transforms:
# for more options see common.factories.transforms_factory.py
- SegRandomFlip:
prob: 0.5
- SegRescale: # consider removing this step
long_size: 608
- SegRandomRescale:
scales: [ 0.5, 2.0 ]
- SegPadShortToCropSize:
crop_size: 512
- SegCropImageAndMask:
crop_size: 512
mode: random
- SegStandardize:
max_value: 255
- SegNormalize:
mean: [0.485, 0.456, 0.406]
std: [0.229, 0.224, 0.225]
- SegConvertToTensor:
mask_output_dtype: long
val_dataset_params:
root_dir: /data/coco/
list_file: instances_val2017.json
samples_sub_directory: images/val2017
targets_sub_directory: annotations
dataset_classes_inclusion_tuples_list:
_target_: super_gradients.training.utils.segmentation_utils.coco_sub_classes_inclusion_tuples_list
cache_labels: False
cache_images: False
transforms:
- SegRescale:
short_size: 512
- SegCropImageAndMask:
crop_size: 512
mode: center
- SegStandardize:
max_value: 255
- SegNormalize:
mean: [0.485, 0.456, 0.406]
std: [0.229, 0.224, 0.225]
- SegConvertToTensor:
mask_output_dtype: long
train_dataloader_params:
shuffle: True
batch_size: 8
num_workers: 8
drop_last: True # drop the last incomplete batch, if dataset size is not divisible by the batch size
val_dataloader_params:
batch_size: 24
num_workers: 8
drop_last: False
---
Src/Super Gradients/Recipes/Dataset Params/Crowdpose Yolo Nas Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/crowdpose_yolo_nas_dataset_params.yaml)
num_joints: 14
OKs sigma values taken from
https://github.com/Jeff-sjtu/CrowdPose/blob/master/crowdpose-api/PythonAPI/crowdposetools/cocoeval.py#L223
oks_sigmas: [0.079, 0.079, 0.072, 0.072, 0.062, 0.062, 0.107, 0.107, 0.087, 0.087, 0.089, 0.089, 0.079, 0.079]
flip_indexes: [ 1, 0, 3, 2, 5, 4, 7, 6, 9, 8, 11, 10, 12, 13]
edge_colors:
- [214, 39, 40] # left_shoulder -> neck
- [148, 103, 189] # right_shoulder -> neck
- [44, 160, 44] # neck -> head
- [188, 189, 34] # left_shoulder -> left_elbow
- [31, 119, 180] # left_elbow -> left_wrist
- [23, 190, 207] # right_shoulder -> right_elbow
- [255, 127, 14] # right_elbow -> right_wrist
- [140, 86, 75] # left_shoulder -> left_hip
- [227, 119, 194] # right_shoulder -> right_hip
- [255, 127, 14] # left_hip -> left_knee
- [31, 119, 180] # left_knee -> left_ankle
- [214, 39, 40] # right_hip -> right_knee
- [44, 160, 44] # right_knee -> right_ankle
- [148, 103, 189] # left_hip -> right_hip
edge_links:
- [0, 13] # left_shoulder -> neck
- [1, 13] # right_shoulder -> neck
- [13, 12] # neck -> head
- [0, 2] # left_shoulder -> left_elbow
- [2, 4] # left_elbow -> left_wrist
- [1, 3] # right_shoulder -> right_elbow
- [3, 5] # right_elbow -> right_wrist
- [0, 6] # left_shoulder -> left_hip
- [1, 7] # right_shoulder -> right_hip
- [6, 8] # left_hip -> left_knee
- [8, 10] # left_knee -> left_ankle
- [7, 9] # right_hip -> right_knee
- [9, 11] # right_knee -> right_ankle
- [6, 7] # left_hip -> right_hip
keypoint_colors:
- [148, 103, 189] # left_shoulder (0)
- [31, 119, 180] # right_shoulder (1)
- [148, 103, 189] # left_elbow (2)
- [31, 119, 180] # right_elbow (3)
- [148, 103, 189] # left_wrist (4)
- [31, 119, 180] # right_wrist (5)
- [148, 103, 189] # left_hip (6)
- [31, 119, 180] # right_hip (7)
- [148, 103, 189] # left_knee (8)
- [31, 119, 180] # right_knee (9)
- [148, 103, 189] # left_ankle (10)
- [31, 119, 180] # right_ankle (11)
- [148, 103, 189] # head (12)
- [31, 119, 180] # neck (13)
image_size: 640
dataset_params_suffix: "default_${dataset_params.image_size}"
train_dataset_params:
data_dir: /data/crowdpose
images_dir: images
json_file: crowdpose_trainval.json
include_empty_samples: True
crowd_annotations_action: mask_as_normal
edge_links: ${dataset_params.edge_links}
edge_colors: ${dataset_params.edge_colors}
keypoint_colors: ${dataset_params.keypoint_colors}
transforms:
- KeypointsRandomHorizontalFlip:
flip_index: ${dataset_params.flip_indexes}
prob: 0.5
- KeypointsBrightnessContrast:
brightness_range: [ 0.8, 1.2 ]
contrast_range: [ 0.8, 1.2 ]
prob: 0.5
- KeypointsHSV:
hgain: 20
sgain: 20
vgain: 20
prob: 0.5
- KeypointsRandomAffineTransform:
max_rotation: 0
min_scale: 0.66
max_scale: 1.5
max_translate: 0.1
image_pad_value: 127
mask_pad_value: 1
prob: 0.75
interpolation_mode: [0, 1, 2, 3, 4]
- KeypointsLongestMaxSize:
max_height: ${dataset_params.image_size}
max_width: ${dataset_params.image_size}
- KeypointsPadIfNeeded:
min_height: ${dataset_params.image_size}
min_width: ${dataset_params.image_size}
image_pad_value: [127, 127, 127]
mask_pad_value: 1
padding_mode: center
- KeypointsImageStandardize:
max_value: 255
- KeypointsRemoveSmallObjects:
min_instance_area: 1
min_visible_keypoints: 1
val_dataset_params:
data_dir: /data/crowdpose
images_dir: images
json_file: crowdpose_test.json
include_empty_samples: True
crowd_annotations_action: no_action
edge_links: ${dataset_params.edge_links}
edge_colors: ${dataset_params.edge_colors}
keypoint_colors: ${dataset_params.keypoint_colors}
transforms:
- KeypointsLongestMaxSize:
max_height: ${dataset_params.image_size}
max_width: ${dataset_params.image_size}
- KeypointsPadIfNeeded:
min_height: ${dataset_params.image_size}
min_width: ${dataset_params.image_size}
image_pad_value: 127
mask_pad_value: 1
padding_mode: bottom_right
- KeypointsImageStandardize:
max_value: 255
train_dataloader_params:
dataset: COCOPoseEstimationDataset
batch_size: 8
num_workers: 8
drop_last: True
pin_memory: False
shuffle: True
collate_fn: YoloNASPoseCollateFN
val_dataloader_params:
dataset: COCOPoseEstimationDataset
batch_size: 24
num_workers: 8
drop_last: False
shuffle: False
pin_memory: False
collate_fn: YoloNASPoseCollateFN
---
Src/Super Gradients/Recipes/Dataset Params/Imagenet Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/imagenet_dataset_params.yaml)
Base recipe for ImageNet Datasets amd Dataloaders.
img_mean: [0.485, 0.456, 0.406] # mean for normalization
img_std: [0.229, 0.224, 0.225] # std for normalization
train_dataset_params:
root: /data/Imagenet/train
transforms:
- RandomResizedCropAndInterpolation:
size: 224
interpolation: default
- RandomHorizontalFlip
- ToTensor
- Normalize:
mean: ${dataset_params.img_mean}
std: ${dataset_params.img_std}
val_dataset_params:
root: /data/Imagenet/val
transforms:
- Resize:
size: 256
- CenterCrop:
size: 224
- ToTensor
- Normalize:
mean: ${dataset_params.img_mean}
std: ${dataset_params.img_std}
train_dataloader_params:
shuffle: True
batch_size: 64
num_workers: 8
drop_last: False
pin_memory: True
val_dataloader_params:
batch_size: 200
num_workers: 8
drop_last: False
pin_memory: True
_convert_: all
---
Src/Super Gradients/Recipes/Dataset Params/Imagenet Efficientnet Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/imagenet_efficientnet_dataset_params.yaml)
defaults:
- imagenet_dataset_params
train_dataset_params:
root: /data/Imagenet/train
transforms:
- RandomResizedCropAndInterpolation:
size: 224
interpolation: random
- RandomHorizontalFlip
- RandAugmentTransform:
config_str: rand-m9-mstd0.5
crop_size: 224
img_mean: ${dataset_params.img_mean} # Use default value from imagenet_dataset_params
- ToTensor
- Normalize:
mean: ${dataset_params.img_mean} # Use default value from imagenet_dataset_params
std: ${dataset_params.img_std} # Use default value from imagenet_dataset_params
- RandomErase:
probability: 0.2
value: random
train_dataloader_params:
batch_size: 64
---
Src/Super Gradients/Recipes/Dataset Params/Imagenet Mobilenetv2 Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/imagenet_mobilenetv2_dataset_params.yaml)
defaults:
- imagenet_dataset_params
train_dataset_params:
root: /data/Imagenet/train
transforms:
- RandomResizedCropAndInterpolation:
size: 224
interpolation: random
- RandomHorizontalFlip
- RandAugmentTransform:
config_str: rand-m9-mstd0.5
crop_size: 224
img_mean: ${dataset_params.img_mean} # Use default value from imagenet_dataset_params
- ToTensor
- Normalize:
mean: ${dataset_params.img_mean} # Use default value from imagenet_dataset_params
std: ${dataset_params.img_std} # Use default value from imagenet_dataset_params
- RandomErase:
probability: 0.2
value: random
train_dataloader_params:
drop_last: True
batch_size: 256
val_dataloader_params:
batch_size: 256
---
Src/Super Gradients/Recipes/Dataset Params/Imagenet Mobilenetv3 Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/imagenet_mobilenetv3_dataset_params.yaml)
defaults:
- imagenet_dataset_params
train_dataloader_params:
batch_size: 128
num_workers: 16
val_dataloader_params:
num_workers: 16
---
Src/Super Gradients/Recipes/Dataset Params/Imagenet RegnetY Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/imagenet_regnetY_dataset_params.yaml)
defaults:
- imagenet_dataset_params
train_dataset_params:
root: /data/Imagenet/train
transforms:
- RandomResizedCropAndInterpolation:
size: 224
interpolation: random
- RandomHorizontalFlip
- RandAugmentTransform:
config_str: rand-m9-mstd0.5
crop_size: 224
img_mean: ${dataset_params.img_mean} # Use default value from imagenet_dataset_params
- ToTensor
- Normalize:
mean: ${dataset_params.img_mean} # Use default value from imagenet_dataset_params
std: ${dataset_params.img_std} # Use default value from imagenet_dataset_params
- RandomErase:
probability: 0.2
value: random
train_dataloader_params:
batch_size: 256
---
Src/Super Gradients/Recipes/Dataset Params/Imagenet Resnet50 Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/imagenet_resnet50_dataset_params.yaml)
defaults:
- imagenet_dataset_params
train_dataset_params:
root: /data/Imagenet/train
transforms:
- RandomResizedCropAndInterpolation:
size: 224
interpolation: random
- RandomHorizontalFlip
- RandAugmentTransform:
config_str: rand-m7-mstd0.5
crop_size: 224
img_mean: ${dataset_params.img_mean} # Use default value from imagenet_dataset_params
- ToTensor
- Normalize:
mean: ${dataset_params.img_mean} # Use default value from imagenet_dataset_params
std: ${dataset_params.img_std} # Use default value from imagenet_dataset_params
val_dataset_params:
root: /data/Imagenet/val
transforms:
- Resize:
size: 236
- CenterCrop:
size: 224
- ToTensor
- Normalize:
mean: ${dataset_params.img_mean}
std: ${dataset_params.img_std}
train_dataloader_params:
batch_size: 236
collate_fn:
_target_: super_gradients.training.datasets.mixup.CollateMixup
mixup_alpha: 0.2
cutmix_alpha: 1.0
label_smoothing: 0.1
---
Src/Super Gradients/Recipes/Dataset Params/Imagenet Resnet50 Kd Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/imagenet_resnet50_kd_dataset_params.yaml)
defaults:
- imagenet_dataset_params
train_transform_args:
interpolation: random
color_jitter: [0.4, 0.4, 0.4]
random_erase_prob: 0.
random_erase_value: random
auto_augment_config_string: rand-m7-mstd0.5
train_dataset_params:
root: /data/Imagenet/train
transforms:
- RandomResizedCropAndInterpolation:
size: 224
interpolation: random
- RandomHorizontalFlip
- RandAugmentTransform:
config_str: rand-m7-mstd0.5
crop_size: 224
img_mean: ${dataset_params.img_mean} # Use default value from imagenet_dataset_params
- ToTensor
- Normalize:
mean: ${dataset_params.img_mean} # Use default value from imagenet_dataset_params
std: ${dataset_params.img_std} # Use default value from imagenet_dataset_params
train_dataloader_params:
batch_size: 192
collate_fn:
_target_: super_gradients.training.datasets.mixup.CollateMixup
mixup_alpha: 0.2
cutmix_alpha: 1.0
label_smoothing: 0.1
sampler:
RepeatAugSampler:
num_repeats: 3
val_dataloader_params:
batch_size: 256
---
Src/Super Gradients/Recipes/Dataset Params/Imagenet Vit Base Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/imagenet_vit_base_dataset_params.yaml)
defaults:
- imagenet_dataset_params
train_dataset_params:
root: /data/Imagenet/train
transforms:
- RandomResizedCropAndInterpolation:
size: 224
interpolation: random
- RandomHorizontalFlip
- RandAugmentTransform:
config_str: rand-m7-mstd0.5
crop_size: 224
img_mean: [0.5, 0.5, 0.5]
- ToTensor
- Normalize:
mean: [0.5, 0.5, 0.5]
std: [0.5, 0.5, 0.5]
val_dataset_params:
root: /data/Imagenet/val
transforms:
- Resize:
size: 249
- CenterCrop:
size: 224
- ToTensor
- Normalize:
mean: [0.5, 0.5, 0.5]
std: [0.5, 0.5, 0.5]
train_dataloader_params:
batch_size: 64
collate_fn:
_target_: super_gradients.training.datasets.mixup.CollateMixup
mixup_alpha: 0.2
cutmix_alpha: 1.0
label_smoothing: 0.1
---
Src/Super Gradients/Recipes/Dataset Params/Mapillary Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/mapillary_dataset_params.yaml)
config_version: '1.2'
config_ignore_values:
'1.2': 65
'2.0': 123
mask_fill_value: ${getitem:${dataset_params.config_ignore_values},${dataset_params.config_version}}
train_dataset_params:
root_dir: /data/mapillary-vistas-dataset_public_v2.0
config_file: config_v${..config_version}.json
samples_sub_directory: training/images
targets_sub_directory: training/v${..config_version}/labels
cache_labels: False
cache_images: False
transforms:
- SegRescale:
long_size: 2048
- SegColorJitter:
brightness: 0.5
contrast: 0.5
saturation: 0.5
- SegRandomFlip:
prob: 0.5
- SegRandomRescale:
scales: [ 0.5, 2.0 ]
- SegPadShortToCropSize:
crop_size: 1024
fill_mask: ${dataset_params.mask_fill_value}
- SegCropImageAndMask:
crop_size: 1024
mode: random
- SegStandardize:
max_value: 255
- SegNormalize:
mean: [0.485, 0.456, 0.406]
std: [0.229, 0.224, 0.225]
- SegConvertToTensor:
mask_output_dtype: long
val_dataset_params:
root_dir: /data/mapillary-vistas-dataset_public_v2.0
config_file: config_v${..config_version}.json
samples_sub_directory: validation/images
targets_sub_directory: validation/v${..config_version}/labels
cache_labels: False
cache_images: False
transforms:
- SegRescale:
long_size: 2048
- SegPadToDivisible:
divisible_value: 32
fill_mask: ${dataset_params.mask_fill_value}
- SegStandardize:
max_value: 255
- SegNormalize:
mean: [0.485, 0.456, 0.406]
std: [0.229, 0.224, 0.225]
- SegConvertToTensor:
mask_output_dtype: long
train_dataloader_params:
shuffle: True
batch_size: 4
num_workers: 8
drop_last: True
pin_memory: True
val_dataloader_params:
# Mapillary validation set include various image sizes.
# It is recommended to Rescale the long size to 2048 then perform validation.
# Unless the default transformation hasn't modified, it is not possible to batch the images to a common size.
batch_size: 1
num_workers: 8
drop_last: False
pin_memory: True
---
Src/Super Gradients/Recipes/Dataset Params/Pascal Aug Segmentation Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/pascal_aug_segmentation_dataset_params.yaml)
defaults:
- pascal_voc_segmentation_dataset_params
- _self_
train_dataset_params:
list_file:
samples_sub_directory:
targets_sub_directory:
---
Src/Super Gradients/Recipes/Dataset Params/Pascal Voc Detection Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/pascal_voc_detection_dataset_params.yaml)
train_dataset_params:
data_dir: ./data/pascal_voc/
input_dim: [320, 320]
transforms:
- DetectionPaddedRescale:
input_dim: ${dataset_params.train_dataset_params.input_dim}
- DetectionTargetsFormatTransform:
input_dim: ${dataset_params.train_dataset_params.input_dim}
output_format: LABEL_CXCYWH
images_dir: images
labels_dir: labels
class_inclusion_list:
max_num_samples:
download: True
val_dataset_params:
data_dir: ./data/pascal_voc/
input_dim: [320, 320]
transforms:
- DetectionPaddedRescale:
input_dim: ${dataset_params.train_dataset_params.input_dim}
- DetectionTargetsFormatTransform:
input_dim: ${dataset_params.train_dataset_params.input_dim}
output_format: LABEL_CXCYWH
images_dir: images/test2007/
labels_dir: labels/test2007/
class_inclusion_list:
max_num_samples:
download: True
train_dataloader_params:
shuffle: True
batch_size: 16
num_workers: 8
drop_last: True
pin_memory: True
worker_init_fn:
_target_: super_gradients.training.utils.utils.load_func
dotpath: super_gradients.training.datasets.datasets_utils.worker_init_reset_seed
collate_fn: DetectionCollateFN
val_dataloader_params:
batch_size: 64
num_workers: 8
drop_last: False
pin_memory: True
collate_fn: DetectionCollateFN
_convert_: all
---
Src/Super Gradients/Recipes/Dataset Params/Pascal Voc Segmentation Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/pascal_voc_segmentation_dataset_params.yaml)
train_dataset_params:
root: /data/pascal_voc_2012
list_file: VOCdevkit/VOC2012/ImageSets/Segmentation/train.txt
samples_sub_directory: VOCdevkit/VOC2012/JPEGImages
targets_sub_directory: VOCdevkit/VOC2012/SegmentationClass
cache_labels: False
cache_images: False
transforms:
# for more options see common.factories.transforms_factory.py
- SegRescale:
long_size: 512
- SegRandomFlip:
prob: 0.5
- SegColorJitter:
brightness: 0.5
contrast: 0.5
saturation: 0.5
- SegRandomRescale:
scales: [ 0.5, 2.0 ]
- SegPadShortToCropSize:
crop_size: 512
fill_mask: 21
- SegCropImageAndMask:
crop_size: 512
mode: random
- SegStandardize:
max_value: 255
- SegNormalize:
mean: [0.485, 0.456, 0.406]
std: [0.229, 0.224, 0.225]
- SegConvertToTensor:
mask_output_dtype: long
val_dataset_params:
root: /data/pascal_voc_2012
list_file: VOCdevkit/VOC2012/ImageSets/Segmentation/val.txt
samples_sub_directory: VOCdevkit/VOC2012/JPEGImages
targets_sub_directory: VOCdevkit/VOC2012/SegmentationClass
cache_labels: False
cache_images: False
transforms:
- SegRescale:
long_size: 512
- SegPadShortToCropSize:
crop_size: 512
fill_mask: 21
- SegCropImageAndMask:
crop_size: 512
mode: center
- SegStandardize:
max_value: 255
- SegNormalize:
mean: [0.485, 0.456, 0.406]
std: [0.229, 0.224, 0.225]
- SegConvertToTensor:
mask_output_dtype: long
train_dataloader_params:
shuffle: True
batch_size: 16
num_workers: 8
drop_last: True
pin_memory: True
val_dataloader_params:
batch_size: 16
num_workers: 8
drop_last: False
pin_memory: True
---
Src/Super Gradients/Recipes/Dataset Params/Roboflow Detection Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/roboflow_detection_dataset_params.yaml)
data_dir: /data/rf100
dataset_name: # Set the name of the dataset you want to use (e.g. "digits-t2eg6")
train_dataset_params:
data_dir: ${..data_dir} # root path to Robflow datasets
dataset_name: ${..dataset_name}
split: train
input_dim: [640, 640]
cache_annotations: True
ignore_empty_annotations: False
transforms:
- DetectionMosaic:
input_dim: ${dataset_params.train_dataset_params.input_dim}
prob: 1.
- DetectionRandomAffine:
degrees: 0. # rotation degrees, randomly sampled from [-degrees, degrees]
translate: 0.1 # image translation fraction
scales: [ 0.5, 1.5 ] # random rescale range (keeps size by padding/cropping) after mosaic transform.
shear: 0.0 # shear degrees, randomly sampled from [-degrees, degrees]
target_size: ${dataset_params.train_dataset_params.input_dim}
filter_box_candidates: False # whether to filter out transformed bboxes by edge size, area ratio, and aspect ratio.
wh_thr: 2 # edge size threshold when filter_box_candidates = True (pixels)
area_thr: 0.1 # threshold for area ratio between original image and the transformed one, when filter_box_candidates = True
ar_thr: 20 # aspect ratio threshold when filter_box_candidates = True
border_value: 128
- DetectionMixup:
input_dim: ${dataset_params.train_dataset_params.input_dim}
mixup_scale: [ 0.5, 1.5 ] # random rescale range for the additional sample in mixup
prob: 1.0 # probability to apply per-sample mixup
flip_prob: 0.5 # probability to apply horizontal flip
- DetectionHSV:
prob: 1.0 # probability to apply HSV transform
hgain: 5 # HSV transform hue gain (randomly sampled from [-hgain, hgain])
sgain: 30 # HSV transform saturation gain (randomly sampled from [-sgain, sgain])
vgain: 30 # HSV transform value gain (randomly sampled from [-vgain, vgain])
- DetectionHorizontalFlip:
prob: 0.5 # probability to apply horizontal flip
- DetectionPaddedRescale:
input_dim: ${dataset_params.train_dataset_params.input_dim}
- DetectionStandardize:
max_value: 255.
- DetectionTargetsFormatTransform:
input_dim: ${dataset_params.train_dataset_params.input_dim}
output_format: LABEL_CXCYWH
class_inclusion_list:
max_num_samples:
with_crowd: False
verbose: 0
train_dataloader_params:
shuffle: True
batch_size: 16
min_samples: 512
num_workers: 4
drop_last: False
pin_memory: True
worker_init_fn:
_target_: super_gradients.training.utils.utils.load_func
dotpath: super_gradients.training.datasets.datasets_utils.worker_init_reset_seed
collate_fn: DetectionCollateFN
val_dataset_params:
data_dir: ${..data_dir} # root path to Robflow datasets
dataset_name: ${..dataset_name}
split: valid
input_dim: [640, 640]
cache_annotations: True
ignore_empty_annotations: False
transforms:
- DetectionPaddedRescale:
input_dim: ${dataset_params.val_dataset_params.input_dim}
pad_value: 114
- DetectionStandardize:
max_value: 255.
- DetectionTargetsFormatTransform:
input_dim: ${dataset_params.val_dataset_params.input_dim}
output_format: LABEL_CXCYWH
class_inclusion_list:
max_num_samples:
with_crowd: True
verbose: 0
val_dataloader_params:
batch_size: 32
num_workers: 4
drop_last: False
shuffle: False
pin_memory: True
collate_fn: CrowdDetectionCollateFN
_convert_: all
---
Src/Super Gradients/Recipes/Dataset Params/Supervisely Persons Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/supervisely_persons_dataset_params.yaml)
dataset_dir: /data/supervisely-persons
batch_size: 8
train_dataset_params:
root_dir: ${..dataset_dir}
list_file: train.csv
cache_labels: False
cache_images: False
transforms:
- SegRandomRescale:
scales: [ 0.25, 1. ]
- SegColorJitter:
brightness: 0.5
contrast: 0.5
saturation: 0.5
- SegRandomFlip:
prob: 0.5
- SegPadShortToCropSize:
crop_size: [ 320, 480 ]
fill_mask: 0
- SegCropImageAndMask:
crop_size: [ 320, 480 ]
mode: random
- SegStandardize:
max_value: 255
- SegNormalize:
mean: [0.485, 0.456, 0.406]
std: [0.229, 0.224, 0.225]
- SegConvertToTensor:
mask_output_dtype: long
val_dataset_params:
root_dir: ${..dataset_dir}
list_file: val.csv
cache_labels: False
cache_images: False
transforms:
- SegResize:
h: 480
w: 320
- SegStandardize:
max_value: 255
- SegNormalize:
mean: [0.485, 0.456, 0.406]
std: [0.229, 0.224, 0.225]
- SegConvertToTensor:
mask_output_dtype: long
train_dataloader_params:
dataset: SuperviselyPersonsDataset
shuffle: True
batch_size: ${..batch_size}
drop_last: True
val_dataloader_params:
dataset: SuperviselyPersonsDataset
batch_size: ${..batch_size}
drop_last: False
_convert_: all
---
Src/Super Gradients/Recipes/Dataset Params/Tiny Imagenet Dataset Params.Yaml (src/super_gradients/recipes/dataset_params/tiny_imagenet_dataset_params.yaml)
defaults:
- imagenet_dataset_params
train_dataset_params:
root: /data/TinyImagenet/train
transforms:
- RandomResizedCropAndInterpolation:
size: 56
- RandomHorizontalFlip
- ToTensor
- Normalize:
mean: [0.4802, 0.4481, 0.3975]
std: [0.2770, 0.2691, 0.2821]
val_dataset_params:
root: /data/TinyImagenet/val
transforms:
- Resize:
size: 64
- CenterCrop:
size: 56
- ToTensor
- Normalize:
mean: [0.4802, 0.4481, 0.3975]
std: [0.2770, 0.2691, 0.2821]
_convert_: all
---
Src/Super Gradients/Recipes/Quantization Params/Default Quantization Params.Yaml (src/super_gradients/recipes/quantization_params/default_quantization_params.yaml)
ptq_only: False # whether to launch QAT, or leave PTQ only
selective_quantizer_params:
calibrator_w: "max" # calibrator type for weights, acceptable types are ["max", "histogram"]
calibrator_i: "histogram" # calibrator type for inputs acceptable types are ["max", "histogram"]
per_channel: True # per-channel quantization of weights, activations stay per-tensor by default
learn_amax: False # enable learnable amax in all TensorQuantizers using straight-through estimator
skip_modules: # optional list of module names (strings) to skip from quantization
calib_params:
histogram_calib_method: "percentile" # calibration method for all "histogram" calibrators, acceptable types are ["percentile", "entropy", "mse"], "max" calibrators always use "max"
percentile: 99.99 # percentile for all histogram calibrators with method "percentile", other calibrators are not affected
num_calib_batches: 16 # number of batches to use for calibration, if None, 512 / batch_size will be used
verbose: False # if calibrator should be verbose
---
Src/Super Gradients/Recipes/Training Hyperparams/Cifar10 Resnet Train Params.Yaml (src/super_gradients/recipes/training_hyperparams/cifar10_resnet_train_params.yaml)
defaults:
- default_train_params
max_epochs: 250
lr_updates:
_target_: numpy.arange
start: 100
stop: 250
step: 50
lr_decay_factor: 0.1
lr_mode: StepLRScheduler
lr_warmup_epochs: 0
initial_lr: 0.1
loss: CrossEntropyLoss
optimizer: SGD
criterion_params: {}
optimizer_params:
weight_decay: 1e-4
momentum: 0.9
metric_to_watch: Accuracy
greater_metric_to_watch_is_better: True
train_metrics_list: # metrics for evaluation
- Accuracy
- Top5
valid_metrics_list: # metrics for evaluation
- Accuracy
- Top5
_convert_: all
---
Src/Super Gradients/Recipes/Training Hyperparams/Cityscapes Default Train Params.Yaml (src/super_gradients/recipes/training_hyperparams/cityscapes_default_train_params.yaml)
defaults:
- default_train_params
max_epochs: 800
lr_mode: PolyLRScheduler
initial_lr: 0.01 # for effective batch_size=32
lr_warmup_epochs: 10
multiply_head_lr: 10.
optimizer: SGD
optimizer_params:
momentum: 0.9
weight_decay: 5e-4
ema: True
ema_params:
decay: 0.9999
beta: 15
decay_type: exp
train_metrics_list:
- PixelAccuracy:
ignore_label: 19
- IoU:
num_classes: 20
ignore_index: 19
valid_metrics_list:
- PixelAccuracy:
ignore_label: 19
- IoU:
num_classes: 20
ignore_index: 19
zero_weight_decay_on_bias_and_bn: True
average_best_models: True
mixed_precision: False
metric_to_watch: IoU
greater_metric_to_watch_is_better: True
_convert_: all
---
Src/Super Gradients/Recipes/Training Hyperparams/Coco2017 Dekr Pose Train Params.Yaml (src/super_gradients/recipes/training_hyperparams/coco2017_dekr_pose_train_params.yaml)
defaults:
- default_train_params
ema: False
ema_params:
decay: 0.9997
decay_type: exp
beta: 20
max_epochs: 150
lr_mode: CosineLRScheduler
cosine_final_lr_ratio: 0.1
batch_accumulate: 1
initial_lr: 1e-3
loss: DEKRLoss
criterion_params:
heatmap_loss: qfl
heatmap_loss_factor: 1.0
offset_loss_factor: 0.1
mixed_precision: True
optimizer: AdamW
optimizer_params:
weight_decay: 0.0001
lr_warmup_steps: 256
warmup_initial_lr: 1e-06
valid_metrics_list:
- PoseEstimationMetrics:
num_joints: ${dataset_params.num_joints}
oks_sigmas: ${dataset_params.oks_sigmas}
max_objects_per_image: 30
post_prediction_callback:
_target_: super_gradients.training.utils.pose_estimation.DEKRPoseEstimationDecodeCallback
max_num_people: 30
keypoint_threshold: 0.05
nms_threshold: 0.05
nms_num_threshold: 8
output_stride: 4
apply_sigmoid: True
phase_callbacks: []
Note: You can uncomment following block to enable visualization of intermediate results during training.
When enabled, these callbacks will save first batch from training & validation to Tensorboard.
This is helpful for debugging and doing visual checks whether predictions are reasonable and transforms are
working as expected.
The only downside is that it tend to bloat Tensorboard logs (Up to ten Gigs for long training regimes).
phase_callbacks:
- DEKRVisualizationCallback:
phase:
_target_: super_gradients.training.utils.callbacks.callbacks.Phase
value: TRAIN_BATCH_END
prefix: "train_"
mean: [ 0.485, 0.456, 0.406 ]
std: [ 0.229, 0.224, 0.225 ]
apply_sigmoid: True
#
- DEKRVisualizationCallback:
phase:
_target_: super_gradients.training.utils.callbacks.callbacks.Phase
value: VALIDATION_BATCH_END
prefix: "val_"
mean: [ 0.485, 0.456, 0.406 ]
std: [ 0.229, 0.224, 0.225 ]
apply_sigmoid: True
metric_to_watch: 'AP'
greater_metric_to_watch_is_better: True
---
Src/Super Gradients/Recipes/Training Hyperparams/Coco2017 Ppyoloe Train Params.Yaml (src/super_gradients/recipes/training_hyperparams/coco2017_ppyoloe_train_params.yaml)
defaults:
- default_train_params
max_epochs: 500
static_assigner_end_epoch: 150
warmup_mode: LinearBatchLRWarmup
warmup_initial_lr: 1e-6
lr_warmup_steps: 1000
lr_warmup_epochs: 0
initial_lr: 2e-3
lr_mode: CosineLRScheduler
cosine_final_lr_ratio: 0.1
zero_weight_decay_on_bias_and_bn: False
batch_accumulate: 1
save_ckpt_epoch_list: [200, 250, 300, 350, 400, 450]
loss: PPYoloELoss
criterion_params:
num_classes: ${arch_params.num_classes}
optimizer: AdamW
optimizer_params:
weight_decay: 0.0001
ema: True
ema_params:
decay: 0.9997
decay_type: threshold
mixed_precision: False
sync_bn: True
valid_metrics_list:
- DetectionMetrics:
score_thres: 0.1
top_k_predictions: 300
num_cls: ${arch_params.num_classes}
normalize_targets: True
post_prediction_callback:
_target_: super_gradients.training.models.detection_models.pp_yolo_e.PPYoloEPostPredictionCallback
score_threshold: 0.01
nms_top_k: 1000
max_predictions: 300
nms_threshold: 0.7
pre_prediction_callback:
phase_callbacks:
- PPYoloETrainingStageSwitchCallback:
static_assigner_end_epoch: ${training_hyperparams.static_assigner_end_epoch}
metric_to_watch: '[email protected]:0.95'
greater_metric_to_watch_is_better: True
_convert_: all
---
Src/Super Gradients/Recipes/Training Hyperparams/Coco2017 Rescoring Train Params.Yaml (src/super_gradients/recipes/training_hyperparams/coco2017_rescoring_train_params.yaml)
defaults:
- default_train_params
ema: True
ema_params:
decay: 0.9997
decay_type: exp
beta: 20
max_epochs: 50
lr_mode: CosineLRScheduler
cosine_final_lr_ratio: 0.1
batch_accumulate: 1
initial_lr: 0.001
loss: RescoringLoss
criterion_params: {}
mixed_precision: False
optimizer: AdamW
optimizer_params:
weight_decay: 0.0001
lr_warmup_steps: 256
warmup_initial_lr: 1e-06
valid_metrics_list:
- PoseEstimationMetrics:
num_joints: ${dataset_params.num_joints}
oks_sigmas: ${dataset_params.oks_sigmas}
max_objects_per_image: 30
post_prediction_callback:
_target_: super_gradients.training.utils.pose_estimation.RescoringPoseEstimationDecodeCallback
apply_sigmoid: True
metric_to_watch: 'AP'
greater_metric_to_watch_is_better: True
---
Src/Super Gradients/Recipes/Training Hyperparams/Coco2017 Ssd Lite Mobilenet V2 Train Params.Yaml (src/super_gradients/recipes/training_hyperparams/coco2017_ssd_lite_mobilenet_v2_train_params.yaml)
defaults:
- default_train_params
ema: True
max_epochs: 400
lr_mode: CosineLRScheduler
cosine_final_lr_ratio: 0.01
batch_accumulate: 1
initial_lr: 0.01
loss: SSDLoss
criterion_params:
alpha: 1.0
dboxes: # OVERRIDEN IN MAIN RECIPE YAML FILE ONCE DBOXES ARE CHOSEN.
optimizer: SGD
optimizer_params:
momentum: 0.9
weight_decay: 0.0005
nesterov: True
lr_warmup_epochs: 3
warmup_momentum: 0.8
warmup_initial_lr: 1e-06
warmup_bias_lr: 0.1
valid_metrics_list:
- DetectionMetrics:
post_prediction_callback:
_target_: super_gradients.training.utils.ssd_utils.SSDPostPredictCallback
conf: 0.001
iou: 0.6
num_cls: 80
metric_to_watch: '[email protected]:0.95'
greater_metric_to_watch_is_better: True
---
Src/Super Gradients/Recipes/Training Hyperparams/Coco2017 Yolo Nas Pose Train Params.Yaml (src/super_gradients/recipes/training_hyperparams/coco2017_yolo_nas_pose_train_params.yaml)
defaults:
- default_train_params
warmup_mode: "LinearBatchLRWarmup"
warmup_initial_lr: 1e-6
lr_warmup_steps: 128
lr_warmup_epochs: 10
initial_lr: 2e-3
lr_mode: cosine
cosine_final_lr_ratio: 0.05
max_epochs: 1000
zero_weight_decay_on_bias_and_bn: True
batch_accumulate: 1
average_best_models: True
save_ckpt_epoch_list: [300, 500]
loss: yolo_nas_pose_loss
criterion_params:
oks_sigmas: ${dataset_params.oks_sigmas}
classification_loss_weight: 1.0
classification_loss_type: focal
regression_iou_loss_type: ciou
iou_loss_weight: 2.5
dfl_loss_weight: 0.01
pose_cls_loss_weight: 1.0
pose_reg_loss_weight: 34.0
pose_classification_loss_type: focal
rescale_pose_loss_with_assigned_score: True
assigner_multiply_by_pose_oks: True
optimizer: AdamW
optimizer_params:
weight_decay: 0.000001
ema: True
ema_params:
decay: 0.997
decay_type: threshold
mixed_precision: True
sync_bn: False
valid_metrics_list:
- PoseEstimationMetrics:
num_joints: ${dataset_params.num_joints}
oks_sigmas: ${dataset_params.oks_sigmas}
max_objects_per_image: 30
post_prediction_callback:
_target_: super_gradients.training.models.pose_estimation_models.yolo_nas_pose.YoloNASPosePostPredictionCallback
pose_confidence_threshold: 0.01
nms_iou_threshold: 0.7
pre_nms_max_predictions: 300
post_nms_max_predictions: 30
phase_callbacks:
# You can uncomment this callback to visualize predictions during training
# - ExtremeBatchPoseEstimationVisualizationCallback:
# keypoint_colors: ${dataset_params.keypoint_colors}
# edge_colors: ${dataset_params.edge_colors}
# edge_links: ${dataset_params.edge_links}
# loss_to_monitor: YoloNASPoseLoss/loss
# max: True
# freq: 1
# max_images: 16
# enable_on_train_loader: True
# enable_on_valid_loader: True
# post_prediction_callback:
# _target_: super_gradients.training.models.pose_estimation_models.yolo_nas_pose.YoloNASPosePostPredictionCallback
# pose_confidence_threshold: 0.1
# nms_iou_threshold: 0.7
# pre_nms_max_predictions: 300
# post_nms_max_predictions: 30
- EarlyStop:
phase:
_target_: super_gradients.training.utils.callbacks.base_callbacks.Phase
value: VALIDATION_EPOCH_END
monitor: AP
mode: max
min_delta: 0.0001
patience: 100
verbose: True
pre_prediction_callback:
metric_to_watch: 'AP'
greater_metric_to_watch_is_better: True
_convert_: all
---
Src/Super Gradients/Recipes/Training Hyperparams/Coco2017 Yolo Nas Train Params.Yaml (src/super_gradients/recipes/training_hyperparams/coco2017_yolo_nas_train_params.yaml)
defaults:
- default_train_params
max_epochs: 300
warmup_mode: LinearBatchLRWarmup
warmup_initial_lr: 1e-6
lr_warmup_steps: 1000
lr_warmup_epochs: 0
initial_lr: 2e-4
lr_mode: CosineLRScheduler
cosine_final_lr_ratio: 0.1
zero_weight_decay_on_bias_and_bn: True
batch_accumulate: 1
save_ckpt_epoch_list: [100, 200, 250]
loss: PPYoloELoss
criterion_params:
use_static_assigner: False
num_classes: ${arch_params.num_classes}
optimizer: AdamW
optimizer_params:
weight_decay: 0.00001
ema: True
ema_params:
decay: 0.9997
decay_type: threshold
mixed_precision: False
sync_bn: True
This is how you can enable visualization of predictions during training
A batch with the largest loss will be visualized for train and valid loaders
Visualization images will be logged using configured logger
phase_callbacks:
- ExtremeBatchDetectionVisualizationCallback:
loss_to_monitor: "PPYoloELoss/loss"
max: True
enable_on_train_loader: False
enable_on_valid_loader: True
classes: ${dataset_params.class_names}
normalize_targets: True
post_prediction_callback:
_target_: super_gradients.training.models.detection_models.pp_yolo_e.PPYoloEPostPredictionCallback
score_threshold: 0.25
nms_top_k: 300
max_predictions: 30
nms_threshold: 0.7
valid_metrics_list:
- DetectionMetrics:
score_thres: 0.1
top_k_predictions: 300
num_cls: ${arch_params.num_classes}
normalize_targets: True
post_prediction_callback:
_target_: super_gradients.training.models.detection_models.pp_yolo_e.PPYoloEPostPredictionCallback
score_threshold: 0.01
nms_top_k: 1000
max_predictions: 300
nms_threshold: 0.7
pre_prediction_callback:
metric_to_watch: '[email protected]:0.95'
greater_metric_to_watch_is_better: True
_convert_: all
---
Src/Super Gradients/Recipes/Training Hyperparams/Coco2017 Yolox Train Params.Yaml (src/super_gradients/recipes/training_hyperparams/coco2017_yolox_train_params.yaml)
defaults:
- default_train_params
max_epochs: 300
lr_mode: CosineLRScheduler
cosine_final_lr_ratio: 0.05
lr_warmup_epochs: 5
lr_cooldown_epochs: 15
initial_lr: 0.02
zero_weight_decay_on_bias_and_bn: True
batch_accumulate: 1
save_ckpt_epoch_list: [285]
loss: YoloXDetectionLoss
criterion_params:
strides: [8, 16, 32] # output strides of all yolo outputs
num_classes: 80
optimizer: SGD
optimizer_params:
momentum: 0.9
weight_decay: 0.0005
nesterov: True
ema: True
mixed_precision: True
valid_metrics_list:
- DetectionMetrics:
normalize_targets: True
post_prediction_callback:
_target_: super_gradients.training.models.detection_models.yolo_base.YoloXPostPredictionCallback
iou: 0.65
conf: 0.01
num_cls: 80
pre_prediction_callback:
phase_callbacks:
- YoloXTrainingStageSwitchCallback:
next_stage_start_epoch: 285
metric_to_watch: '[email protected]:0.95'
greater_metric_to_watch_is_better: True
_convert_: all
---
Src/Super Gradients/Recipes/Training Hyperparams/Coco Segmentation Shelfnet Lw Train Params.Yaml (src/super_gradients/recipes/training_hyperparams/coco_segmentation_shelfnet_lw_train_params.yaml)
defaults:
- default_train_params
max_epochs: 150
initial_lr: 5e-3
loss: ShelfNetOHEMLoss
optimizer: SGD
mixed_precision: True
batch_accumulate: 3
lr_mode: PolyLRScheduler
optimizer_params:
momentum: 0.9
weight_decay: 1e-4
nesterov: False
load_opt_params: False
train_metrics_list:
- PixelAccuracy
- IoU:
num_classes: 21
valid_metrics_list:
- PixelAccuracy
- IoU:
num_classes: 21
metric_to_watch: IoU
greater_metric_to_watch_is_better: True
---
Src/Super Gradients/Recipes/Training Hyperparams/Default Train Params.Yaml (src/super_gradients/recipes/training_hyperparams/default_train_params.yaml)
resume: False # Whether to continue training from ckpt from the latest run, within the same experiment name.
run_id: # ID of run to resume from the same experiment.
resume_path: # Explicit checkpoint path (.pth file) to use to resume training.
resume_from_remote_sg_logger: False # bool (default=False), When true, ckpt_name (checkpoint filename
to resume i.e ckpt_latest.pth bydefault) will be downloaded into the experiment checkpoints directory
prior to loading weights, then training is resumed from that checkpoint. The source is unique to
every logger, and currently supported for WandB loggers only.
#
IMPORTANT: Only works for experiments that were ran with sg_logger_params.save_checkpoints_remote=True.
IMPORTANT: For WandB loggers, one must also pass the run id through the wandb_id arg in sg_logger_params.
ckpt_name: ckpt_latest.pth # The checkpoint (.pth file) filename in CKPT_ROOT_DIR/EXPERIMENT_NAME/ to use when resume=True and resume_path=None
lr_mode: # Union[str, Mapping]
# when str: Learning rate scheduling policy, one of ["StepLRScheduler", "PolyLRScheduler", "CosineLRScheduler", "ExponentialLRScheduler", "FunctionLRScheduler"]
# when Mapping: refers to a torch.optim.lr_scheduler._LRScheduler, following the below API: lr_mode = {LR_SCHEDULER_CLASS_NAME: {LR_SCHEDULER_KWARGS, "phase": XXX, "metric_name": XXX)
lr_schedule_function: # Learning rate scheduling function to be used when lr_mode is 'FunctionLRScheduler'.
lr_warmup_epochs: 0 # number of epochs for learning rate warm up - see https://arxiv.org/pdf/1706.02677.pdf (Section 2.2).
lr_warmup_steps: 0 # number of warmup steps (Used when warmup_mode=LinearBatchLRWarmup)
lr_cooldown_epochs: 0 # epochs to cooldown LR (i.e the last epoch from scheduling view point=max_epochs-cooldown)
warmup_initial_lr: # Initial lr for LinearEpochLRWarmup/LinearBatchLRWarmup. When none is given, initial_lr/(warmup_epochs+1) will be used.
step_lr_update_freq: # (float) update frequency in epoch units for computing lr_updates when lr_mode=StepLRScheduler.
cosine_final_lr_ratio: 0.01 # final learning rate ratio (only relevant when lr_mode='CosineLRScheduler')
warmup_mode: LinearEpochLRWarmup # learning rate warmup scheme, currently ['LinearEpochLRWarmup', 'LinearEpochLRWarmup', 'LinearBatchLRWarmup'] are supported
lr_updates:
_target_: super_gradients.training.utils.utils.empty_list # This is a workaround to instantiate a list using _target_. If we would instantiate as "lr_updates: []",
# we would get an error every time we would want to overwrite lr_updates with a numpy array.
pre_prediction_callback: # callback modifying images and targets right before forward pass.
optimizer: SGD # Optimization algorithm. One of ['Adam','SGD','RMSProp'] corresponding to the torch.optim optimizers
optimizer_params: {} # when optimizer is one of ['Adam','SGD','RMSProp'], it will be initialized with optimizer_params.
load_opt_params: True # Whether to load the optimizers parameters as well when loading a model's checkpoint
zero_weight_decay_on_bias_and_bn: False # whether to apply weight decay on batch normalization parameters or not
loss: # Loss function for training (str as one of SuperGradient's built in options, or torch.nn.module)
criterion_params: {} # when loss is one of SuperGradient's built in options, it will be initialized with criterion_params.
ema: False # whether to use Model Exponential Moving Average
ema_params: # parameters for the ema model.
decay: 0.9999
decay_type: exp
beta: 15
train_metrics_list: [] # Metrics to log during training. For more information on torchmetrics see https://torchmetrics.rtfd.io/en/latest/.
valid_metrics_list: [] # Metrics to log during validation. For more information on torchmetrics see https://torchmetrics.rtfd.io/en/latest/
metric_to_watch: Accuracy # will be the metric which the model checkpoint will be saved according to
greater_metric_to_watch_is_better: True # When choosing a model's checkpoint to be saved, the best achieved model is the one that maximizes the metric_to_watch when this parameter is set to True
launch_tensorboard: False # Whether to launch a TensorBoard process.
tensorboard_port: # port for tensorboard process
tb_files_user_prompt: False # Asks User for Tensorboard Deletion Prompt
save_tensorboard_to_s3: False # whether to save tb to s3
precise_bn: False # Whether to use precise_bn calculation during the training.
precise_bn_batch_size: # the effective batch size we want to calculate the batchnorm on.
sync_bn: False # Whether to convert bn layers to synched bn (for accurate stats in DDP).
silent_mode: False # Silents the Print outs
mixed_precision: False # Whether to use mixed precision or not.
save_ckpt_epoch_list: [] # indices where the ckpt will save automatically
average_best_models: True # If set, a snapshot dictionary file and the average model will be saved
dataset_statistics: False # add a dataset statistical analysis and sample images to tensorboard
batch_accumulate: 1 # number of batches to accumulate before every backward pass
run_validation_freq: 1 # The frequency in which validation is performed during training.
run_test_freq: 1 # The frequency in which test is performed during training.
save_model: True # Whether to save the model checkpoints
seed: 42 # seed for reproducibility
phase_callbacks: [] # list of callbacks to be applied at specific phases.
log_installed_packages: True # when set, the list of all installed packages (and their versions) will be written to the tensorboard
clip_grad_norm : # Defines a maximal L2 norm of the gradients. Values which exceed the given value will be clipped
ckpt_best_name: ckpt_best.pth
max_train_batches: # For debug- when not None- will break out of inner train loop
(i.e iterating over train_loader) when reaching this number of batches.
max_valid_batches: # For debug- when not None- will break out of inner valid loop
(i.e iterating over valid_loader) when reaching this number of batches.
sg_logger: base_sg_logger
sg_logger_params:
tb_files_user_prompt: False # Asks User for Tensorboard Deletion Prompt
launch_tensorboard: False
tensorboard_port:
save_checkpoints_remote: False # upload checkpoint files to s3
save_tensorboard_remote: False # upload tensorboard files to s3
save_logs_remote: False # upload log files to s3
monitor_system: True # Monitor and write to tensorboard the system statistics, such as CPU usage, GPU, ...
torch_compile: False # Enable or disable use of torch.compile to optimize the model (Requires Pytorch 2.0)
torch_compile_loss: False # Enable or disable use of torch.compile to optimize the loss (Requires Pytorch 2.0)
torch.compile options from https://pytorch.org/docs/stable/generated/torch.compile.html
torch_compile_options:
mode: reduce-overhead # default / reduce-overhead / max-autotune
fullgraph: False # Whether it is ok to break model into several subgraphs
dynamic: False # Use dynamic shape tracing
backend: inductor # backend to be used
options: # A dictionary of options to pass to the backend.
disable: False # Turn torch.compile() into a no-op for testing
finetune: False # Whether to freeze a fixed part of the model. Supported only for models that implement get_finetune_lr_dict.
# The model's class method get_finetune_lr_dict should return a dictionary, mapping lr to the
# unfrozen part of the network, in the same fashion as using initial_lr.
_convert_: all
---
Src/Super Gradients/Recipes/Training Hyperparams/Imagenet Efficientnet Train Params.Yaml (src/super_gradients/recipes/training_hyperparams/imagenet_efficientnet_train_params.yaml)
defaults:
- default_train_params
max_epochs: 450
lr_mode: StepLRScheduler
step_lr_update_freq: 2.4
initial_lr: 0.016
lr_warmup_epochs: 3
warmup_initial_lr: 1e-6
lr_decay_factor: 0.97
optimizer: RMSpropTF
optimizer_params:
momentum: 0.9
weight_decay: 1e-5
eps: 0.001
ema: True
ema_params:
decay: 0.9999
decay_type: constant
loss: CrossEntropyLoss
criterion_params:
smooth_eps: 0.1
metric_to_watch: Accuracy
greater_metric_to_watch_is_better: True
save_ckpt_epoch_list: [50, 100, 150, 200]
average_best_models: True
mixed_precision: True
zero_weight_decay_on_bias_and_bn: True
train_metrics_list: # metrics for evaluation
- Accuracy
- Top5
valid_metrics_list: # metrics for evaluation
- Accuracy
- Top5
_convert_: all
---
Src/Super Gradients/Recipes/Training Hyperparams/Imagenet Mobilenetv2 Train Params.Yaml (src/super_gradients/recipes/training_hyperparams/imagenet_mobilenetv2_train_params.yaml)
defaults:
- default_train_params
max_epochs: 450
lr_mode: StepLRScheduler
initial_lr: 0.032 # for total batch-size of 512
lr_decay_factor: 0.973
lr_updates:
_target_: numpy.arange
start: 2.4
stop: 450
step: 2.4
lr_warmup_epochs: 5
optimizer: RMSpropTF
optimizer_params:
weight_decay: 0.00001
momentum: 0.9
alpha: 0.9
eps: 0.001
loss: CrossEntropyLoss
zero_weight_decay_on_bias_and_bn: True
ema: True
ema_params:
decay: 0.9999
mixed_precision: True
metric_to_watch: Accuracy
greater_metric_to_watch_is_better: True
train_metrics_list: # metrics for evaluation
- Accuracy
- Top5
valid_metrics_list: # metrics for evaluation
- Accuracy
- Top5
_convert_: all
---
Src/Super Gradients/Recipes/Training Hyperparams/Imagenet Mobilenetv3 Train Params.Yaml (src/super_gradients/recipes/training_hyperparams/imagenet_mobilenetv3_train_params.yaml)
defaults:
- default_train_params
max_epochs: 150
lr_mode: CosineLRScheduler
initial_lr: 0.1
optimizer: SGD
optimizer_params:
weight_decay: 0.00004
lr_warmup_epochs: 5
loss: CrossEntropyLoss
criterion_params:
smooth_eps: 0.1
zero_weight_decay_on_bias_and_bn: True
ema: True
metric_to_watch: Accuracy
greater_metric_to_watch_is_better: True
train_metrics_list: # metrics for evaluation
- Accuracy
- Top5
valid_metrics_list: # metrics for evaluation
- Accuracy
- Top5
_convert_: all
---
Src/Super Gradients/Recipes/Training Hyperparams/Imagenet RegnetY Train Params.Yaml (src/super_gradients/recipes/training_hyperparams/imagenet_regnetY_train_params.yaml)
defaults:
- default_train_params
max_epochs: 450
lr_mode: StepLRScheduler
step_lr_update_freq: 2.4
initial_lr: 0.016
lr_warmup_epochs: 3
warmup_initial_lr: 1e-6
lr_decay_factor: 0.97
optimizer: RMSpropTF
optimizer_params:
momentum: 0.9
weight_decay: 1e-5
eps: 0.001
ema: True
ema_params:
decay_type: constant
decay: 0.9999
loss: CrossEntropyLoss
criterion_params:
smooth_eps: 0.1
metric_to_watch: Accuracy
greater_metric_to_watch_is_better: True
save_ckpt_epoch_list: [50, 100, 150, 200]
average_best_models: True
mixed_precision: True
zero_weight_decay_on_bias_and_bn: True
train_metrics_list: # metrics for evaluation
- Accuracy
- Top5
valid_metrics_list: # metrics for evaluation
- Accuracy
- Top5
_convert_: all
---
Src/Super Gradients/Recipes/Training Hyperparams/Imagenet Repvgg Train Params.Yaml (src/super_gradients/recipes/training_hyperparams/imagenet_repvgg_train_params.yaml)
defaults:
- default_train_params
max_epochs: 120
lr_mode: CosineLRScheduler
initial_lr: 0.1
cosine_final_lr_ratio: 0
loss: CrossEntropyLoss
zero_weight_decay_on_bias_and_bn: True
average_best_models: True
metric_to_watch: Accuracy
greater_metric_to_watch_is_better: True
train_metrics_list: # metrics for evaluation
- Accuracy
- Top5
valid_metrics_list: # metrics for evaluation
- Accuracy
- Top5
_convert_: all
---
Src/Super Gradients/Recipes/Training Hyperparams/Imagenet Resnet50 Kd Train Params.Yaml (src/super_gradients/recipes/training_hyperparams/imagenet_resnet50_kd_train_params.yaml)
defaults:
- default_train_params
max_epochs: 610
initial_lr: 5e-3
lr_mode: CosineLRScheduler
lr_warmup_epochs: 5
lr_cooldown_epochs: 10
ema: True
mixed_precision: True
zero_weight_decay_on_bias_and_bn: True
optimizer: Lamb
optimizer_params:
weight_decay: 0.02
loss: CrossEntropyLoss
train_metrics_list: # metrics for evaluation
- Accuracy
- Top5
valid_metrics_list: # metrics for evaluation
- Accuracy
- Top5
_convert_: all
---
Src/Super Gradients/Recipes/Training Hyperparams/Imagenet Resnet50 Train Params.Yaml (src/super_gradients/recipes/training_hyperparams/imagenet_resnet50_train_params.yaml)
defaults:
- default_train_params
max_epochs: 400
initial_lr: 0.1
lr_mode: CosineLRScheduler
lr_warmup_epochs: 5
ema: False
save_ckpt_epoch_list: [ 50, 100, 150, 200, 300 ]
mixed_precision: True
zero_weight_decay_on_bias_and_bn: True
loss: CrossEntropyLoss
train_metrics_list: # metrics for evaluation
- Accuracy
- Top5
valid_metrics_list: # metrics for evaluation
- Accuracy
- Top5
metric_to_watch: Accuracy
greater_metric_to_watch_is_better: True
_convert_: all
---
Src/Super Gradients/Recipes/Training Hyperparams/Imagenet Vit Train Params.Yaml (src/super_gradients/recipes/training_hyperparams/imagenet_vit_train_params.yaml)
defaults:
- default_train_params
max_epochs: 10
initial_lr: 0.03
lr_mode: CosineLRScheduler
cosine_final_lr_ratio: 0
lr_warmup_epochs: 1
warmup_initial_lr: 0.
warmup_mode: LinearEpochLRWarmup
ema: False
loss: CrossEntropyLoss
clip_grad_norm: 1
optimizer: SGD
optimizer_params:
weight_decay: 0
momentum: 0.9
train_metrics_list: # metrics for evaluation
- Accuracy
- Top5
valid_metrics_list: # metrics for evaluation
- Accuracy
- Top5
metric_to_watch: Accuracy
greater_metric_to_watch_is_better: True
average_best_models: False
_convert_: all
---
Src/Super Gradients/Recipes/Training Hyperparams/Supervisely Default Train Params.Yaml (src/super_gradients/recipes/training_hyperparams/supervisely_default_train_params.yaml)
defaults:
- default_train_params
max_epochs: 100
lr_mode: CosineLRScheduler
cosine_final_lr_ratio: 0.01
initial_lr: 0.1
lr_warmup_epochs: 0
multiply_head_lr: 1.
optimizer: SGD
optimizer_params:
momentum: 0.9
weight_decay: 5e-4
ema: True
ema_params:
decay: 0.999
decay_type: exp
beta: 15
train_metrics_list:
- BinaryIOU
valid_metrics_list:
- BinaryIOU
zero_weight_decay_on_bias_and_bn: True
average_best_models: True
mixed_precision: False
metric_to_watch: target_IOU
greater_metric_to_watch_is_better: True
---
Src/Super Gradients/Training/Datasets/Dataset Setup Instructions (src/super_gradients/training/datasets/Dataset_Setup_Instructions.md)
Computer Vision Datasets Setup
SuperGradients provides multiple Datasets implementations.
Classification Datasets
<details>
<summary>Cifar10</summary>
Supports download
from super_gradients.training.datasets import Cifar10
dataset = Cifar10(..., download=True)</details>
<details>
<summary>Imagenet</summary>
1. Download imagenet dataset:
- https://image-net.org/download.php
2. Unzip:
Imagenet
βββtrain
β βββn02093991
β β βββn02093991_1001.JPEG
β β βββn02093991_1004.JPEG
β β βββ...
β βββn02093992
β βββ...
βββval
βββn02093991
βββn02093992
βββ...3. Instantiate the dataset:
from super_gradients.training.datasets import ImageNetDataset
train_set = ImageNetDataset(root='.../Imagenet/train', ...)
valid_set = ImageNetDataset(root='.../Imagenet/val', ...)</details>
Detection Datasets
<details>
<summary>Coco</summary>
1. Download coco dataset:
- annotations: http://images.cocodataset.org/annotations/annotations_trainval2017.zip
- train2017: http://images.cocodataset.org/zips/train2017.zip
- val2017: http://images.cocodataset.org/zips/val2017.zip
2. Unzip and organize it as below:
coco
βββ annotations
β ββ instances_train2017.json
β ββ instances_val2017.json
β ββ ...
βββ images
βββ train2017
β ββ 000000000001.jpg
β ββ ...
βββ val2017
ββ ...3. Instantiate the dataset:
from super_gradients.training.datasets import COCODetectionDataset
train_set = COCODetectionDataset(data_dir='.../coco', subdir='images/train2017', json_file='instances_train2017.json', ...)
valid_set = COCODetectionDataset(data_dir='.../coco', subdir='images/val2017', json_file='instances_val2017.json', ...)</details>
<details>
<summary>PascalVOC 2007 & 2012</summary>
Supports download
from super_gradients.training.datasets import PascalVOCDetectionDataset
train_set = PascalVOCDetectionDataset(download=True, ...)Dataset Structure:
Dataset structure:
ββimages
β ββ train2012
β ββ val2012
β ββ VOCdevkit
β β ββ VOC2007
β β β βββJPEGImages
β β β βββSegmentationClass
β β β βββImageSets
β β β βββImageSets/Segmentation
β β β βββImageSets/Main
β β β βββImageSets/Layout
β β β βββAnnotations
β β β βββSegmentationObject
β β βββVOC2012
β β βββJPEGImages
β β βββSegmentationClass
β β βββImageSets
β β βββImageSets/Segmentation
β β βββImageSets/Main
β β βββImageSets/Action
β β βββImageSets/Layout
β β βββAnnotations
β β βββSegmentationObject
β ββtrain2007
β ββtest2007
β ββval2007
ββlabels
ββtrain2012
ββval2012
ββtrain2007
ββtest2007
ββval2007</details>
<details>
<summary>Yolo/Darknet format</summary>
1. Download your dataset (can be from https://roboflow.com/universe)
2. You should have a structure similar to this.
data_dir
βββ train/test/val
βββ images
β ββ 0001.jpg
β ββ 0002.jpg
β ββ ...
βββ labels
ββ 0001.txt
ββ 0002.txt
ββ ...Note: train/test/val folders are not required, any folder structure is supported.
3. Instantiate the dataset:
from super_gradients.training.datasets import YoloDarknetFormatDetectionDataset
data_set = YoloDarknetFormatDetectionDataset(data_dir='<path-to>/data_dir', images_dir="<train/test/val>/images", labels_dir="<train/test/val>/labels", classes=["<to-fill>"])</details>
Segmentation Datasets
<details>
<summary>Cityscapes</summary>
1. Download dataset:
- a. Cityscapes dataset:
- gtFine: https://www.cityscapes-dataset.com/file-handling/?packageID=1
- leftImg8bit: https://www.cityscapes-dataset.com/file-handling/?packageID=3
- b. metadata folder: https://deci-pretrained-models.s3.amazonaws.com/cityscape_lists.zip
2. a. Unzip and organize cityscapes dataset as below:
root_dir (in recipe default to /data/cityscapes)
ββββ gtFine
β βββ test
β β βββ berlin
β β β βββ berlin_000000_000019_gtFine_color.png
β β β βββ berlin_000000_000019_gtFine_instanceIds.png
β β β βββ ...
β β βββ bielefeld
β β β βββ ...
β β βββ ...
β ββββ train
β β βββ ...
β ββββ val
β βββ ...
ββββ leftImg8bit
βββ test
β βββ ...
ββββ train
β βββ ...
ββββ val
βββ ...2. b. Unzip and organize metadata folder as below:
lists
βββ labels.csv
βββ test.lst
βββ train.lst
βββ trainval.lst
βββ val.lst
βββ auto_labelling.lst2. c. Move Metadata folder to the Cityscapes folder
root_dir (in recipe default to /data/cityscapes)
ββββ gtFine
β βββ ...
ββββ leftImg8bit
β βββ ...
ββββ lists
βββ ...3. Instantiate the dataset:
from super_gradients.training.datasets import CityscapesDatasettrain_set = CityscapesDataset(root_dir='.../root_dir', list_file='lists/train.lst', labels_csv_path='lists/labels.csv', ...)
4. AutoLabelling dataset [Optional]
Cityscapes AutoLabelled dataset were introduced by NVIDIA research group
in the paper:
"Hierarchical Multi-Scale Attention for Semantic Segmentation".
AutoLabelled refer to the refinement of the Cityscapes coarse data and pseudo
labels generation using their suggested Hierarchical multi-scale attention model.
* To download the AutoLabelled labels please refer to the original
repo.
Unzip and rename the folder to AutoLabelling as described bellow.
* Download the coarse RGB images from cityscapes official site,
leftImg8bit_train_extra: https://www.cityscapes-dataset.com/file-handling/?packageID=4
root_dir (in recipe default to /data/cityscapes)
ββββ gtFine
β βββ test
β β βββ ...
β ββββ train
β β βββ ...
β ββββ val
β βββ ...
ββββ leftImg8bit
β βββ test
β β βββ ...
β ββββ train
β β βββ ...
β ββββ val
β βββ ...
ββββ AutoLabelling
β ββββ train_extra
β βββ ...
ββββ leftImg8bit
ββββ train_extra
βββ ...</details>
<details>
<summary>Coco</summary>
1. Download coco dataset:
- annotations: http://images.cocodataset.org/annotations/annotations_trainval2017.zip
- train2017: http://images.cocodataset.org/zips/train2017.zip
- val2017: http://images.cocodataset.org/zips/val2017.zip
2. Unzip and organize it as below:
coco
βββ annotations
β ββ instances_train2017.json
β ββ instances_val2017.json
β ββ ...
βββ images
βββ train2017
β ββ 000000000001.jpg
β ββ ...
βββ val2017
ββ ...
3. Instantiate the dataset:
from super_gradients.training.datasets import CoCoSegmentationDataSet
train_set = CoCoSegmentationDataSet(data_dir='.../coco', subdir='images/train2017', json_file='instances_train2017.json', ...)
valid_set = CoCoSegmentationDataSet(data_dir='.../coco', subdir='images/val2017', json_file='instances_val2017.json', ...)</details>
<details>
<summary>Pascal VOC 2012</summary>
1. Download pascal datasets:
- VOC 2012: http://host.robots.ox.ac.uk/pascal/VOC/voc2012/VOCtrainval_11-May-2012.tar
2. Unzip and organize it as below:
pascal_voc_2012
βββVOCdevkit
βββVOC2012
βββJPEGImages
βββSegmentationClass
βββImageSets
β βββSegmentation
β β βββ train.txt
β βββMain
β βββAction
β βββLayout
βββAnnotations
βββSegmentationObject
3. Instantiate the dataset:
from super_gradients.training.datasets import PascalVOC2012SegmentationDataSettrain_set = PascalVOC2012SegmentationDataSet(
root='.../pascal_voc_2012',
list_file='VOCdevkit/VOC2012/ImageSets/Segmentation/train.txt',
samples_sub_directory='VOCdevkit/VOC2012/JPEGImages',
targets_sub_directory='VOCdevkit/VOC2012/SegmentationClass',
...
)
valid_set = PascalVOC2012SegmentationDataSet(
root='.../pascal_voc_2012',
list_file='VOCdevkit/VOC2012/ImageSets/Segmentation/val.txt',
samples_sub_directory='VOCdevkit/VOC2012/JPEGImages',
targets_sub_directory='VOCdevkit/VOC2012/SegmentationClass',
...
)
</details>
<details>
<summary>Pascal AUG 2012</summary>
1. Download pascal dataset
- AUG 2012: https://www2.eecs.berkeley.edu/Research/Projects/CS/vision/grouping/semantic_contours/benchmark.tgz
2. Unzip and organize it as below:
pascal_voc_2012
βββVOCaug
βββ aug.txt
βββ dataset
βββinst
βββimg
βββcls
3. Instantiate the dataset:
from super_gradients.training.datasets import PascalAUG2012SegmentationDataSettrain_set = PascalAUG2012SegmentationDataSet(
root='.../pascal_voc_2012',
list_file='VOCaug/dataset/aug.txt',
samples_sub_directory='VOCaug/dataset/img',
targets_sub_directory='VOCaug/dataset/cls',
...
)
NOTE: this dataset is only available for training. To test, please use PascalVOC2012SegmentationDataSet.
</details>
<details>
<summary>Pascal AUG & VOC 2012</summary>
1. Download pascal datasets:
- VOC 2012: http://host.robots.ox.ac.uk/pascal/VOC/voc2012/VOCtrainval_11-May-2012.tar
- AUG 2012: https://www2.eecs.berkeley.edu/Research/Projects/CS/vision/grouping/semantic_contours/benchmark.tgz
2. Unzip and organize it as below:
pascal_voc_2012
ββVOCdevkit
β βββVOC2012
β βββJPEGImages
β βββSegmentationClass
β βββImageSets
β β βββSegmentation
β β β βββ train.txt
β β βββMain
β β βββAction
β β βββLayout
β βββAnnotations
β βββSegmentationObject
βββVOCaug
βββ aug.txt
βββ dataset
βββinst
βββimg
βββcls
3. Instantiate the dataset:
from super_gradients.training.datasets import PascalVOCAndAUGUnifiedDataset
train_set = PascalVOCAndAUGUnifiedDataset(root='.../pascal_voc_2012', ...) NOTE: this dataset is only available for training. To test, please use PascalVOC2012SegmentationDataSet.
</details>
<details>
<summary>Supervisely Persons</summary>
1. Download supervisely dataset:
- https://deci-pretrained-models.s3.amazonaws.com/supervisely-persons.zip
2. Unzip:
supervisely-persons
βββimages
β βββimage-name.png
β βββ...
βββimages_600x800
β βββimage-name.png
β βββ...
βββmasks
βββmasks_600x800
3. Instantiate the dataset:
from super_gradients.training.datasets import SuperviselyPersonsDataset
train_set = SuperviselyPersonsDataset(root_dir='.../supervisely-persons', list_file='train.csv', ...)
valid_set = SuperviselyPersonsDataset(root_dir='.../supervisely-persons', list_file='val.csv', ...)NOTE: this dataset is only available for training. To test, please use PascalVOC2012SegmentationDataSet.
</details>
Pose Estimation Datasets
<details>
<summary>COCO 2017</summary>
1. Download coco dataset:
- annotations: http://images.cocodataset.org/annotations/annotations_trainval2017.zip
- train2017: http://images.cocodataset.org/zips/train2017.zip
- val2017: http://images.cocodataset.org/zips/val2017.zip
2. Unzip and organize it as below:
coco
βββ annotations
β ββ person_keypoints_train2017.json
β ββ person_keypoints_val2017.json
β ββ ...
βββ images
βββ train2017
β ββ 000000000001.jpg
β ββ ...
βββ val2017
ββ ...
3. Instantiate the dataset:
from super_gradients.training.datasets import COCOKeypointsDataset
train_set = COCOKeypointsDataset(data_dir='.../coco', images_dir='images/train2017', json_file='annotations/instances_train2017.json', ...)
valid_set = COCOKeypointsDataset(data_dir='.../coco', images_dir='images/val2017', json_file='annotations/instances_val2017.json', ...)</details>
---
Src/Super Gradients/Training/Models/Implemented Model Architectures (src/super_gradients/training/models/Implemented Model Architectures.md)
Implemented Model Architectures
Image Classification
- DensNet (Densely Connected Convolutional Networks) - Densely Connected Convolutional Networks https://arxiv.org/pdf/1608.06993.pdf
- DPN - Dual Path Networks https://arxiv.org/pdf/1707.01629
- EfficientNet - https://arxiv.org/abs/1905.11946
- GoogleNet - https://arxiv.org/pdf/1409.4842
- LeNet - https://yann.lecun.com/exdb/lenet/
- MobileNet - Efficient Convolutional Neural Networks for Mobile Vision Applications https://arxiv.org/pdf/1704.04861
- MobileNet v2 - https://arxiv.org/pdf/1801.04381
- MobileNet v3 - https://arxiv.org/pdf/1905.02244
- PNASNet - Progressive Neural Architecture Search Networks https://arxiv.org/pdf/1712.00559
- Pre-activation ResNet - https://arxiv.org/pdf/1603.05027
- RegNet - https://arxiv.org/pdf/2003.13678.pdf
- RepVGG - Making VGG-style ConvNets Great Again https://arxiv.org/pdf/2101.03697.pdf
- ResNet - Deep Residual Learning for Image Recognition https://arxiv.org/pdf/1512.03385
- ResNeXt - Aggregated Residual Transformations for Deep Neural Networks https://arxiv.org/pdf/1611.05431
- SENet - Squeeze-and-Excitation Networkshttps://arxiv.org/pdf/1709.01507
- ShuffleNet - https://arxiv.org/pdf/1707.01083
- ShuffleNet v2 - Efficient Convolutional Neural Network for Mobile
Deviceshttps://arxiv.org/pdf/1807.11164
- VGG - Very Deep Convolutional Networks for Large-scale Image Recognition https://arxiv.org/pdf/1409.1556
Object Detection
- CSP DarkNet
- DarkNet-53
- SSD (Single Shot Detector) - https://arxiv.org/pdf/1512.02325
- YOLOX - https://arxiv.org/abs/2107.08430
Semantic Segmentation
- PP-LiteSeg - https://arxiv.org/pdf/2204.02681v1.pdf
- DDRNet (Deep Dual-resolution Networks) - https://arxiv.org/pdf/2101.06085.pdf
- LadderNet - Multi-path networks based on U-Net for medical image segmentation https://arxiv.org/pdf/1810.07810
- RegSeg - Rethink Dilated Convolution for Real-time Semantic Segmentation https://arxiv.org/pdf/2111.09957
- ShelfNet - https://arxiv.org/pdf/1811.11254
- STDC - Rethinking BiSeNet For Real-time Semantic Segmentation https://arxiv.org/pdf/2104.13188
</details>
---
Src/Super Gradients/Training/Models/User Models/README (src/super_gradients/training/models/user_models/README.md)
<div "center">
<img src="documentation/assets/SG_img/SG - Horizontal.png" width="600"/>
<br/><br/>
Introduction
This page demonstrates how you can register your own models, so that SuperGradients can access it with a name
str, forexample, when training from a recipe config
architecture: my_custom_model.Usage
1. Create a new Python module in this folder (e.g.
.../user_models/my_model.py).2. Define your PyTorch model (
torch.nn.Module) in the new module.3. Import the
@register decorator from super_gradients.training.models.model_registry import register and apply it to your model.* The decorator can be applied directly to the class or to a function returning the class.
* The decorator takes an optional
name: str argument. If not specified, the decorated class/function name will be registered.Example
import torch.nn as nn
import torch.nn.functional as Ffrom super_gradients.training.utils.registry import register_model
@register_model('my_conv_net') # will be registered as "my_conv_net"
class MyConvNet(nn.Module):
def __init__(self, num_classes):
super().__init__()
self.conv1 = nn.Conv2d(3, 6, 5)
self.pool = nn.MaxPool2d(2, 2)
self.conv2 = nn.Conv2d(6, 16, 5)
self.fc1 = nn.Linear(16 5 5, 120)
self.fc2 = nn.Linear(120, 84)
self.fc3 = nn.Linear(84, num_classes)
def forward(self, x):
x = self.pool(F.relu(self.conv1(x)))
x = self.pool(F.relu(self.conv2(x)))
x = torch.flatten(x, 1)
x = F.relu(self.fc1(x))
x = F.relu(self.fc2(x))
x = self.fc3(x)
return x
or
@register_model()
def myconvnet_for_cifar10(): # will be registered as "myconvnet_for_cifar10"
return MyConvNet(num_classes=10)---
Tests/Unit Tests/Configs/Cifar10 Multiple Test.Yaml (tests/unit_tests/configs/cifar10_multiple_test.yaml)
defaults:
- cifar10_resnet
test_dataloaders:
cifar10: cifar10_val
cifar10_v2: cifar10_val
dataset_params:
train_dataloader_params:
num_workers: 0
val_dataloader_params:
num_workers: 0
test_dataset_params:
cifar10:
root: ./data/cifar10
train: False
transforms:
- Resize:
size: 32
- ToTensor
- Normalize:
mean:
- 0.4914
- 0.4822
- 0.4465
std:
- 0.2023
- 0.1994
- 0.2010
target_transform: null
download: True
cifar10_v2:
root: ./data/cifar10
train: False
transforms:
- Resize:
size: 32
- ToTensor
- Normalize:
mean:
- 0.5
- 0.5
- 0.5
std:
- 0.2
- 0.2
- 0.2
target_transform: null
download: True
hydra:
searchpath:
- pkg://super_gradients.recipes
---
.Github/ISSUE TEMPLATE/Bug Report.Yaml (.github/ISSUE_TEMPLATE/bug_report.yaml)
name: π Bug Report
description: Create a report to help us reproduce and fix the bug
body:
- type: markdown
attributes:
value: >
#### Before submitting a bug, please make sure the issue hasn't been already addressed by searching through the existing and past issues.
#### You can also try using our AI helper to get a fast solution
[](https://docs.deci.ai/super-gradients/latest/documentation/source/welcome.html?autoClick=true)
- type: textarea
attributes:
label: π Describe the bug
description: |
Please provide a clear and concise description of what the bug is.
If relevant, add a minimal example so that we can reproduce the error by running the code. It is very important for the snippet to be as succinct (minimal) as possible, so please take time to trim down any irrelevant code to help us debug efficiently. We are going to copy-paste your code and we expect to get the same result as you did: avoid any external data, and include the relevant imports, etc. For example:
`` python
# All necessary imports at the beginning
from super_gradients.common.object_names import Models
from super_gradients.training import models
# A succinct reproducing example trimmed down to the essential parts:
model = models.get(Models.YOLO_NAS_L, pretrained_weights="coco")
...
Please also paste or describe the results you observe instead of the expected results. If you observe an error, please paste the error message including the full traceback of the exception. It may be relevant to wrap error messages intriple quotes blocks` ``.
placeholder: |
A clear and concise description of what the bug is.
` python
# Sample code to reproduce the problem
text
The error message you got, with the full traceback.
textvalidations:
required: true- type: textarea
attributes:
label: Versions
description: |
Please run the following and paste the output below.
sh
wget https://raw.githubusercontent.com/pytorch/pytorch/main/torch/utils/collect_env.py
# For security purposes, please check the contents of collect_env.py before running it.
python collect_env.py
textvalidations:
required: true- type: markdown
attributes:
value: >
Thanks for contributing π!
---
.Github/ISSUE TEMPLATE/Documentation.Yaml (.github/ISSUE_TEMPLATE/documentation.yaml)
name: π Documentation Improvement
description: Suggest an improvement or provide feedback on the project's documentation
body:
- type: markdown
attributes:
value: >
#### Before submitting a documentation improvement, please make sure it hasn't been already addressed by searching through existing documentation or in an open issues.
- type: textarea
attributes:
label: π Documentation Improvement
description: |
Please provide a clear and concise description of the improvement you'd like to suggest or the feedback you have regarding the project's documentation.
If relevant, you can include specific sections, pages, or examples that need improvement, along with your suggestions or ideas for enhancement.
If applicable, you can include any relevant code snippets.
validations:
required: true
- type: markdown
attributes:
value: >
Thank you for your contribution to improving the documentation! π
---
.Github/ISSUE TEMPLATE/Feature Request.Yaml (.github/ISSUE_TEMPLATE/feature_request.yaml)
name: π Feature Request
description: Suggest a new feature or enhancement for the project
body:
- type: markdown
attributes:
value: >
#### Before submitting a feature request, please make sure it hasn't already been suggested or discussed by searching through existing issues.
- type: textarea
attributes:
label: π Feature Request
description: |
Please provide a clear and concise description of the new feature or enhancement you'd like to suggest for the project.
Explain the problem or need that the feature aims to address. Provide as much detail as possible to help others understand the value and feasibility of the requested feature.
If applicable, you can include code snippets, examples, or any other relevant information to support your feature request.
placeholder: |
Clear and concise description of the new feature.
python
# How you would like to use this feature
textvalidations:
required: true- type: textarea
attributes:
label: Proposed Solution (Optional)
description: |
If you have any ideas or suggestions for how the requested feature could be implemented, you can provide them here. This can include high-level approaches, specific implementation details, or any other relevant information.
placeholder: |
Clear and concise proposed solution.
python
# How you think this feature could be implemented
textvalidations:
required: false- type: markdown
attributes:
value: |
Thank you for suggesting a new feature! Your contribution is appreciated, and we will consider your request. We also encourage you and other community members to actively contribute to the project by addressing the features you suggest.
Feel free to open a pull request and help us bring these ideas to life! If you're new to contributing, check out our contributing guidelines
for guidance on getting started. π
---
.Github/ISSUE TEMPLATE/Question.Yaml (.github/ISSUE_TEMPLATE/question.yaml)
name: π‘ Question
description: Ask a question to get help or clarification
body:
- type: markdown
attributes:
value: >
### Before submitting a question, please make sure it hasn't been already addressed by searching through existing documentation or in ah open issues.
#### You can also try using our AI helper to get a fast answer
[](https://docs.deci.ai/super-gradients/latest/documentation/source/welcome.html?autoClick=true)
- type: textarea
attributes:
label: π‘ Your Question
description: |
Please provide a clear and concise question about the project. Be as specific as possible to facilitate effective responses.
Include any relevant code snippets or examples to support your question. This will help us understand the context of your question better.
placeholder: |
Your clear and concise question here.
python
# Relevant code snippet (if applicable)
textvalidations:
required: true- type: textarea
attributes:
label: Versions
description: |
To help us understand the context better, you can run the following and paste the output below.
sh
wget https://raw.githubusercontent.com/pytorch/pytorch/main/torch/utils/collect_env.py
# For security purposes, please check the contents of collect_env.py before running it.
python collect_env.py
``validations:
required: false
- type: markdown
attributes:
value: >
Thanks for asking your question! Our community will do their best to help you. π
---
.Github/Workflows/Codeql.Yml (.github/workflows/codeql.yml)
name: "CodeQL"
on:
push:
branches: [ 'master' ]
pull_request:
# The branches below must be a subset of the branches above
branches: [ 'master' ]
schedule:
- cron: '39 12 5'
jobs:
analyze:
name: Analyze
runs-on: ${{ (matrix.language == 'swift' && 'macos-latest') || 'ubuntu-latest' }}
timeout-minutes: ${{ (matrix.language == 'swift' && 120) || 360 }}
permissions:
actions: read
contents: read
security-events: write
strategy:
fail-fast: false
matrix:
language: [ 'python' ]
# CodeQL supports [ 'cpp', 'csharp', 'go', 'java', 'javascript', 'python', 'ruby' ]
# Use only 'java' to analyze code written in Java, Kotlin or both
# Use only 'javascript' to analyze code written in JavaScript, TypeScript or both
# Learn more about CodeQL language support at https://aka.ms/codeql-docs/language-support
steps:
- name: Checkout repository
uses: actions/checkout@v3
# Initializes the CodeQL tools for scanning.
- name: Initialize CodeQL
uses: github/codeql-action/init@v2
with:
languages: ${{ matrix.language }}
# If you wish to specify custom queries, you can do so here or in a config file.
# By default, queries listed here will override any specified in a config file.
# Prefix the list here with "+" to use these queries and those in the config file.
# For more details on CodeQL's query packs, refer to: https://docs.github.com/en/code-security/code-scanning/automatically-scanning-your-code-for-vulnerabilities-and-errors/configuring-code-scanning#using-queries-in-ql-packs
# queries: security-extended,security-and-quality
# Autobuild attempts to build any compiled languages (C/C++, C#, Go, Java, or Swift).
# If this step fails, then you should remove it and run the build manually (see below)
- name: Autobuild
uses: github/codeql-action/autobuild@v2
# βΉοΈ Command-line programs to run using the OS shell.
# π See https://docs.github.com/en/actions/using-workflows/workflow-syntax-for-github-actions#jobsjob_idstepsrun
# If the Autobuild fails above, remove it and uncomment the following three lines.
# modify them (or add more) to build your code if your project, please refer to the EXAMPLE below for guidance.
# - run: |
# echo "Run, Build Application using script"
# ./location_of_script_within_repo/buildscript.sh
- name: Perform CodeQL Analysis
uses: github/codeql-action/analyze@v2
with:
category: "/language:${{matrix.language}}"
---
.Github/Workflows/Dependency Review.Yaml (.github/workflows/dependency_review.yaml)
name: 'Dependency Review'
on: [pull_request]
permissions:
contents: read
jobs:
dependency-review:
runs-on: ubuntu-latest
steps:
- name: 'Checkout Repository'
uses: actions/checkout@v3
- name: 'Dependency Review'
uses: actions/dependency-review-action@v1
---
.Github/Workflows/Integration Tests Rc.Yaml (.github/workflows/integration_tests_rc.yaml)
name: 'Release Candidate Integration Tests'
on:
push:
tags:
- '[0-9]+.[0-9]+.[0-9]+rc[0-9]+'
jobs:
release-integration-tests:
runs-on: ubuntu-latest
steps:
- name: Calling CircleCI job
shell: bash
run: |
curl --request POST \
--url https://circleci.com/api/v2/project/gh/Deci-AI/algo-integration-tests/pipeline \
--header 'Circle-Token: ${{ secrets.CIRCLE_CI_TOKEN }}' \
--header 'content-type: application/json' \
--data '{"parameters":{"sg_workflow_sg_version":"${{github.ref_name}}", "sg_workflow_install_from_scratch":false, "sg_workflow_run":true, "sg_workflow_rc":true }}'
---
.Github/Workflows/Integration Tests Release.Yaml (.github/workflows/integration_tests_release.yaml)
name: 'Release Integration Tests'
on:
push:
tags:
- '[0-9]+.[0-9]+.[0-9]+'
jobs:
release-integration-tests:
runs-on: ubuntu-latest
steps:
- name: Calling CircleCI job
shell: bash
run: |
curl --request POST \
--url https://circleci.com/api/v2/project/gh/Deci-AI/algo-integration-tests/pipeline \
--header 'Circle-Token: ${{ secrets.CIRCLE_CI_TOKEN }}' \
--header 'content-type: application/json' \
--data '{"parameters":{"sg_workflow_sg_version":"${{github.ref_name}}", "sg_workflow_install_from_scratch":true, "sg_workflow_run":true, "sg_workflow_rc":false }}'
---