1 · Tracking and re-identification
Detect each person, track them from frame to frame and keep the same identity throughout the sequence.
M1 project · Computer vision
Face recognition assumes that the person cooperates: uncovered face, gaze towards the camera, normal gait. This M1 project, carried out by two students over three months at ISEN Nantes, measures that degree of cooperation automatically from a plain video.
Under controlled conditions, face recognition systems exceed 99% correct verification. In an access-control corridor or a public space, their performance drops as soon as the person does not cooperate: averted gaze, covered face, tilted head, running past.
Assess a person's degree of cooperation automatically, in real time and without disturbing them. The aim is not to replace face recognition, but to give it contextual analysis:
M1 project at ISEN Nantes, lasting three months, carried out as a team of two: we worked together on every task. It was supervised by Khadidja Ould Amer and Cyril Barrelet.
An agile method in five sprints, each delivering a feature that can be tested on its own: state of the art and architecture, face and gaze modules, gait analysis, complementary modules, optimisation and wrap-up. The code lived in a private GitHub repository, with one branch per module.
Run on a standard processor, without a dedicated graphics card, accept a video file or a webcam as input, and keep no biometric data by default.
Four independent blocks, each runnable on its own, written in Python.
Detect each person, track them from frame to frame and keep the same identity throughout the sequence.
Estimate the skeleton and walking speed, and spot running, unusual gait or a person on the floor.
Measure, zone by zone, what hides the face: hand, mask, object.
Estimate head orientation, gaze direction and eye opening to decide whether the person cooperates.
The central block: without a stable identity, no measurement can be followed over time.
YOLOv8m-seg detects people and provides a mask for each; BoT-SORT tracks them from frame to frame. The confidence threshold is lowered to 0.30 so as not to lose people seen in profile or half visible.
InsightFace provides a face embedding, kept only if the face is large enough, well detected and geometrically plausible, which rules out the backs of heads. A body embedding is computed on the person's mask. The two are fused, with the body weighing slightly more: clothing separates two to five people filmed together quite well.
For videos processed afterwards, a second pipeline measures torso colour on each person's mask, recovers the reference colours with K-means, then assigns identities frame by frame with the Hungarian algorithm. A diagnostic file flags ambiguous passages, which a corrections file can settle.
YOLOv8-pose estimates the 17 skeleton keypoints of each person. Speed is computed from the displacement of the hip centre; the scale in metres is derived from the size of the skeleton, set against an average height of 1.70 m, and the measurement is then smoothed.
A majority vote over ten frames keeps the class from flickering. The module was tested on the public Penn Action dataset and on our own videos.
Two blocks sharing the same face landmarks.
MediaPipe FaceMesh places 478 landmarks on the face, which delimit anatomical zones: eyes, nose, mouth, cheeks, forehead. The person's skin colour is calibrated on the first frames, in the Y'CbCr space, which separates brightness from colour. A zone where too few pixels look like skin is considered covered.
The global score, from 0 to 100, weights the zones by their importance for recognition: eyes and nose count three times as much as the forehead. It was tested with a hand in front of the face, a medical mask, an object held up and varied lighting.
Head orientation is estimated by solving PnP from six face points: nose, chin, eye corners and mouth corners. Eye opening is measured with the eye aspect ratio, and gaze direction from the position of the iris.
The person is judged cooperative if both eyes are open, if their gaze is on the camera and if their head is turned by no more than 25° sideways and 20° up or down. A vote over fourteen frames stabilises the display.
The user draws the authorised zones on the image, for instance a pavement. They are converted into a mask, and at every frame the ground point of each tracked person is tested: a person spending more than 30% of the time outside the zone is flagged.
For each identity, the algorithm keeps the most usable frame of the sequence. Its score combines face quality, how frontal the face is and the absence of occlusion.
All six tasks of the subject were completed; the figures below come from the tests in the project report.
Person detection and segmentation, skeleton estimation.
Multi-person tracking, chosen over DeepSORT for its better handling of occlusions.
Face embeddings (ArcFace) for re-identification.
478 face landmarks, including the iris.
Video reading, colour spaces, PnP solving, annotated rendering.
Optimised CPU inference; training on GPU.
The specifications, the project management and the detail of each block.