menu_book

Knowledge Base

Documentation, guides, and resources for Noldus products.

EthoVision XT 19 - Set Up an Experiment - Deep Learning: Basics

Last updated: Jul 28, 2026

Deep Learning: Basics

Background Information

NOTE: The following information only applies to tracking of rodents when using the Deep learning function. See also Adjust the Settings for Nose-Tail Base Detection (Deep Learning).

Deep learning is a type of machine learning. While machine learning is a general category that encompasses all sorts of mathematical tools that help a computer learn by experience, deep learning refers to the use of deep neural networks, where "deep" indicates that the networks are made of multiple, hidden layers of neurons or decision nodes.

In a deep neural network, intermediate layers are placed between the input layer, which receives the data (for example the RGB values of the pixels that make up a bitmap picture), and the output layer, which represents the categories of classification; for example, when classifying a picture, "the picture is of a cat" or "the picture is not of a cat".

Deep Learning In EthoVision XT

A deep neural network can find structures in unstructured data, like pictures and video images. They can recognize recurring patterns, such as the eyes of humans in a number of portraits, and relationships between those patterns. Because of the layered structure of decision nodes, deep neural networks can learn to represent data at various levels, from low levels such as edges, colors and curves, to higher levels such as semicircles (a combination of a curve and a straight edge), squares (a combination of straight edges), up to even higher, more abstract levels (concepts), such as "handwriting", "dark object", "a tail" and so on.

EthoVision XT uses a trained network, that is, the network has learned to extract features from a number of video images of rodents of various colors and in various backgrounds, where the nose and the tail-base were previously annotated. During tracking, the network analyzes a portion of the image that includes the detected subject. It creates a map of probability of occurrence for both the nose and the tail-base points, and finally makes an estimate of the position of the nose point and the tail-base point based on the highest probability.

With deep learning, the detection of the body points is less dependent on the detected contour of the subject. You can see this effect in cases with low contrast between the subject and the background. For example, a dark mouse is only partly detected when it rears with the forepaws placed on the cage wall; however, the neural network can still find the nose point.

Convolutional Networks

EthoVision XT uses a deep Convolutional Neural Network (CNN) to find the nose and tail-base points in each sampled video image. CNNs are particularly suitable to classify images based on spatial relationships. CNNs resemble the structures of the cells in the visual cortex of our brain. The visual cortex has small regions of cells that are sensitive to specific regions of the visual field. Hubel and Wiesel (Journal of Physiology 165: 559–568, 1963) showed that some neurons fired only in the presence of edges of a certain orientation. Some neurons fired when exposed to vertical edges and some when shown horizontal or diagonal edges. The basis of convolutional networks is the idea that specialized components in the network have specific tasks, that is, to look for specific characteristics in the image.

Features Detected With Deep Learning

When you track subjects with deep learning, the main features of the subjects, for example the body center point, are calculated in different ways — either predicted by the neural network or calculated based on the contour of the detected subject. When tracking multiple subjects with deep learning, some of those features are not calculated at all.

One Subject Per Arena

  • Center-point: calculated from contour
  • Nose-point: calculated using deep learning
  • Tail-base point: calculated using deep learning
  • Body elongation: calculated from contour
  • Body area: calculated from contour
  • Body changed area: calculated from contour
  • Head direction line: calculated from contour and nose-point
  • Body elongation is used to calculate the dependent variable Body Elongation and Body Elongation State.
  • Body area and Body changed area are used to calculate Mobility and Mobility State.
  • The Head direction line is used to calculate Head Direction and Head Directed to Zone.

Two Subjects Per Arena

  • Center-point: calculated using deep learning
  • Nose-point: calculated using deep learning
  • Tail-base point: calculated using deep learning
  • Body elongation: not calculated
  • Body area: not calculated
  • Body changed area: not calculated
  • Head direction line: not calculated

Deep Learning: Requirements

Main Topics

  • Graphics card (GPU)
  • Video source
  • Sample rate
  • Subject species, color and size
  • Number of subjects and arenas
  • Video image
  • Video length
  • Individual marking (for two-subject tracking)
  • Test apparatus and background
  • Recording protocol
  • Behavior Recognition
  • Test results

Graphics Card (GPU)

Neural networks make calculations over huge data matrices, and therefore require substantial computation power. In order for the deep learning tracking technique to work in EthoVision XT, you need a Graphics Processing Unit (GPU, or graphics card) that is able to sustain those computations.

Deep learning makes use of the TensorRT software development kit (version 8.6.1.6), which is built on the cuDNN deep neural network library, which in turn relies on the CUDA computing platform. For this reason, the GPU driver must support CUDA runtime version 12.2.

  • More information on graphics cards (GPUs)
  • Install a graphics card for deep learning

Video Source

Interested in this solution?

Leave your details and we will reach out to discuss how we can support your research goals.

Noldus is here to assist you throughout the whole process.

shopping_bag
check_circle

Thank you!

We'll get back to you shortly.

error

Please correct the following errors:

error

error

error

error

By clicking Submit, you consent to Noldus processing your data as described in our privacy policy.