Evaluation
During the evaluation phase, the organizers will collect the annotations and landmark predictions submitted by the participating groups. The analysis will be conducted in two consecutive stages. The first stage will assess the manually generated annotations, whereas the second will focus on the submitted landmark predictions.
Annotation
The annotation analysis will include the subjects, which were annotated by each participating group. The primary objective of this stage is to collect multi-centre large-scale annotations and quantify and characterize inter-annotator variability.
Probability maps will be used to visualize the spatial distribution of the annotations, complemented by descriptive statistics such as the mean and standard deviation. The analysis will also examine whether annotation variability differs between participating centres and whether it is associated with the annotators’ levels of confidence and expertise.
The submitted annotations will be further processed and harmonized to serve as template and training data for Task 2.
Prediction
Since anatomical landmarks are, by their very nature, subjective in the identification, it is difficult to define a “ground truth.” Therefore, predictions for the previously unseen subjects will be assessed using a consensus-based approach. The evaluation will compare the spatial distances among submissions from the participating groups and examine differences between the applied prediction methods.