“Signal Processing in the AI era” was the tagline of this year’s IEEE International Conference on Acoustics, Speech and Signal Processing, taking place in Rhodes, Greece.
In this context, Brent de Weerdt, Xiangyu Yang, Boris Joukovsky, Alex Stergiou and Nikos Deligiannis presented ETRO’s research during poster sessions and oral presentations, with novel ways to process and understand graph, video, and audio data. Nikos Deligiannis chaired a session on Graph Deep Learning, attended the IEEE T-IP Editorial Board Meeting, and had the opportunity to meet with collaborators from the VUB-Duke-Ugent-UCL joint lab.
Featured articles:

On June 20th 2024 at 16:00, Yangxintong Lyu will defend their PhD entitled “DEEP-LEARNING-BASED MULTI-MODAL FUSION FOR TRAFFIC IMAGE DATA PROCESSING”.
Everybody is invited to attend the presentation in room I.0.02, or digitally via this link.
In recent years, deep-learning-based technologies have significantly developed, which is driven by a large amount of data associated with task-specific labels. Among the various formats used for representing object attributes in computer vision, RGB images stand out as a ubiquitous choice. Their value extends to traffic-related applications, particularly in the realms of autonomous driving and intelligent surveillance systems. By using an autonomous driving system, a car is capable of navigating and operating with diminished human interactions, while traffic conditions can be monitored and analysed by an intelligent system. Essentially, the techniques reduce human error and improve road safety, which significantly impacts our daily life.
Although many visual-based traffic analysis tasks can indeed be effectively solved by leveraging features extracted from a sole RGB channel, certain unresolved challenges persist that introduce extra difficulties under certain situations. First of all, extracting complicated information becomes demanding, especially under erratic lighting conditions, raising a need for auxiliary clues. Secondly, obtaining large-scale accurate labels for challenging tasks remains time-consuming, costly, and arduous. The former prompts exploration into capturing and exploiting additional information such that the objects can be observed from diverse aspects; in contrast, the latter requires either an increase in the volume of available data or the capability to learn from other datasets that already possess perfect labels.
In this thesis, we tackle multi-modal data fusion and data scarcity for intelligent transportation systems. Our first contribution is a novel RGB-Thermal fusion neural network for semantic segmentation. It ensures the segmentation under limited illumination. Our second contribution is a 3D-prior-based framework for monocular vehicle 6D pose estimation. The use of 3D geometry avoids the ill-posed pose prediction from a single camera viewpoint. Thanks to the extra 3D information, our novel method can handle distant and occluded vehicles. The third contribution is a real-world, large-scale vehicle make and model dataset that contains the most popular brands operating in Europe. Moreover, we propose a two-branch deep learning vehicle make and model recognition paradigm to reduce inter-make ambiguity. The last contribution is a weakly supervised vehicle 6D pose estimation paradigm by adapting knowledge built based on a novel synthetic dataset. The dataset includes a large amount of accurate labels for vehicles. By learning from the synthetic dataset, our method allows the significant reduction of expensive real-life vehicle pose annotations.
Comprehensive experimental results reveal that the newly introduced datasets hold significant promise for deep-learning-based processing of traffic image data. Moreover, the proposed methods surpass the existing baselines in the literature. Our research not only yields high-quality scientific publications but also underscores its value across both academic and industrial domains.
On October 3rd 2024 at 16:00, Michiel Dhont will defend their PhD entitled “UNSUPERVISED ANALYTICS FOR MULTI-SOURCE TIME SERIES DATA”.
Everybody is invited to attend the presentation in room I.2.02 or online via this link.
There has been an explosion of impressive success stories recently with deep learning (DL) approaches in various fields such as natural language processing, computer vision, healthcare, and robotics. The advent of transformers has further amplified the capabilities of DL models to understand complex patterns, establishing them as a cornerstone of modern AI advancements across a broad spectrum of applications. Initially, transformers revolutionised large language models like GPT-4 and BERT, enabling them to process and generate human-like text with remarkable coherence and accuracy. Now, their impressive performance is also being demonstrated in other domains, extending their impact. Given sufficient high-quality labelled data and computational resources, DL models are able to achieve an accuracy that were previously unattainable.
Unfortunately, most of the real-world application contexts generate datasets which significantly diverge from the idealised benchmark datasets used to validate novel AI methodologies. Real-world data is typically characterised with presence of noise, missing values, complicated parameter names, different data types, lack of ground-truth, context-dependent features, etc. The latter makes it very challenging to immediately dive into any AI model application since it is often not clear which modelling paradigm best suit the problem at hand. This PhD research is built around the conception and validation of a heuristic data analytics methodology with the aim to benefit maximum from the different facets, while mitigating the imperfections, of real-world datasets.
Nowadays, most of the available datasets originating from industrial activities are composed of multitude of different parameters. The inherent multi-source nature of such datasets makes it impossible to directly integrate different data types without information loss. To address this challenge, a multi-view data integration approach has been devised as a part of this PhD, which identifies and considers different data views explicitly, allowing to fully harness the richness of heterogeneous datasets while retaining all relevant information.
The ongoing trend of increasingly more data being captured, goes parallel with an increasing complexity of extracting valuable insights from it. For instance, the remote monitoring of infrastructures (e.g., roads and power supplies) typically generates complex spatio-temporal data streams captured at high sampling rate across different locations. Combining and making sense of such data streams is not trivial. In this PhD research, a spatio-temporal profiling methodology is proposed, allowing to uncover insightful spatial patterns and dependencies while taking full advantage of the temporal dimension. Additionally, the exciting domain of visual analytics has been explored, resulting into the conception of several novel visualisation approaches, blending advanced visualisation with intelligent analysis to effectively reveal key iinsights.
By far, the hardest challenge associated with the analysis of real-world data is the lack of ground truth, which limits the choice of learning paradigms to only unsupervised ones. In this PhD research, a novel modelling framework is conceived, capable of extracting semantically interpretable states from unlabelled data. The latter facilitates a better understanding of system behaviour in terms of state transitions and allows to convert the unsupervised data modelling problem into a supervised one. Several different neural and neuro-symbolic forecasting workflows have been proposed for this purpose
– The language of the programme is English, and the requirements are described here: http://www.vub.ac.be/en/study/applied-sciences-and-engineering-applied-computer-science#admission-criteria;
– If cannot provide a proof of sufficient knowledge of English, then, after the positive evaluation of your academic background and all other criteria to follow the program are met, through an interview the professors will also assess your level of English and a final decision will be taken.
We are very proud of Brent De Weerdt (prom N. Deligiannis), Joris Wuts (prom. J. Vandemeulebroucke) and Silvia Zaccadi (prom. B. Jansen) who received their scholarship from FWO aspirant strategic basic research for the next 2+2 years. Way to go!

Abel DĂaz Berenguer (Cuba) joined ETRO in 2017 and obtained his PhD in 2021. His father was a Civil engineer, and during his childhood, he spent a lot of time with him in construction works. This triggered his curiosity to build and create things. Since Abel was a kid, he wanted to study ¨something¨ with computers and never had any doubt about studying engineering in informatics sciences.
The PhD program made him feel thoroughly responsible for your research project. Advisors progressively introduce you into the research environment and educate you on digging deeper into fundamental theories by promoting critical thinking. Abel noticed a lack of motivation to participate in dissemination activities that promote science communication. Students also did not feel the need to communicate and to encourage more collaboration between other fellows working in the same or different fields.
Abel enjoyed the Writing Bootcamp of the doctoral Training program very much. The three days course about writing scientific articles offered mainly many good tips about approaching the scientific writing process, which was a great help during the PhD.
Abel has great memories of chats, coffees, and plenty of amusing moments with colleagues. Those moments allowed him to overcome challenging days. The collaboration with close colleagues has been outstanding. Research is a process of knowledge sharing and co-creation with excellent colleagues that became family that stood by my side on long working days and nights.
Abel grew during his PhD into a person with more critical thinking towards solving any problem in life. Be enthusiastic and passionate about your research. Enjoy the learning process with perseverance, dedication, engagement, and curiosity for innovating. Push boundaries and never give up share knowledge and be a team player.
Abel wants to land in an academic environment, help others learn and learn from them. In a position to benefit society, share knowledge and contribute to building a better world for our children.