Research Publications

Publications, preprints, and theses.

Download .bib

2026

Context Saturation in Zero-Shot Time-Series Foundation Models

Miguel Nogales, Luca Butera, Alberto Ferrante, Cesare Alippi

ICML 2026 Workshop on Combining Theory and Benchmarks: Towards A Virtuous Cycle to Understand and Guarantee Foundation Model Performance · 2026

Despite time series foundation models (TSFMs) supporting variable input lengths, they are usually evaluated using the longest input possible, depending on data availability and model input capacity. This practice risks conflating different factors impacting performance, and leaves end-users lacking principled guidelines for input length selection. To study the relationship between input length $W$ and model performance, we introduce the context saturation length (CSL): the minimum input length required to achieve a target fraction of a model's peak forecasting performance. For time series with a dominant seasonal period $P$, we show that context saturation can be reliably achieved with at least $W \approx 2P$. This is supported by results on both synthetic and real-world benchmarks, where performance-to-input-length curves align when considering the period-normalized input length $W/P$. Additionally, we demonstrate similar behavior on time series generated by auto-regressive processes, when normalizing the input length by the process memory length. Our findings demonstrate that, for periodic time series where the dominant period can be reliably estimated, the input length can be selected a-priori, avoiding hyper-parameter search, with negligible sacrifice of performance. Moreover, our work suggests that model-agnostic methodologies, based on the inherent characteristics of the target time series, can provide practical guidelines for input length selection, and lead to the design of principled benchmarks to evaluate TSFMs.

Why Do Time Series Models Need Long Context Windows?

Luca Butera, Giovanni De Felice, Andrea Cini, Cesare Alippi

arXiv preprint arXiv:2606.01999 · 2026

Modern deep learning models for forecasting groups of time series rely on increasingly longer observation windows. However, the benefit of increasing the window size is often simply attributed to capturing long-range dependencies, and broader discussion on how global forecasting models leverage input observations has been limited. In this paper, we show that forecasting groups of time series involves two objectives: (i) generative process identification (GPI), i.e., inferring the specific process generating the input sequence, and (ii) conditional forecasting (CF), i.e., predicting future values given input observations. From this perspective, optimal predictions can be interpreted as an average over plausible data-generating processes, weighted by their likelihood given the input window. This suggests another explanation for the benefits of long context windows: they reduce the uncertainty about which specific process is generating the input time series during operation. We prove that even for processes with memory length $P$, an input window size strictly larger than $P$ is necessary to achieve the minimum attainable error. Finally, we show how decoupling GPI and CF can improve computational scalability without compromising accuracy. Experiments on synthetic and real-world data validate our insights and their relevance for designing forecasting architectures.

2025

An Efficient Ground-aerial Transportation System for Pest Control Enabled by AI-based Autonomous Nano-UAVs

Luca Crupi, Luca Butera, Alberto Ferrante, Alessandro Giusti, Daniele Palossi

ACM J. Auton. Transport. Syst. · 2025

Efficient crop production requires early detection of pest outbreaks and timely treatments; we consider a solution based on a fleet of multiple autonomous miniaturized unmanned aerial vehicles (nano-UAVs) to visually detect pests and a single slower heavy vehicle that visits the detected outbreaks to deliver treatments. To cope with the extreme limitations aboard nano-UAVs, e.g., low-resolution sensors and sub-100 mW computational power budget, we design, fine-tune, and optimize a tiny image-based convolutional neural network (CNN) for pest detection. Despite the small size of our CNN (i.e., 0.58 GOps/inference), on our dataset, it scores a mean average precision (mAP) of 0.79 in detecting harmful bugs, i.e., 14\% lower mAP but 32{\texttimes} fewer operations than the best-performing CNN in the literature. Our CNN runs in real-time at 6.8 frame/s, requiring 33 mW on a GWT GAP9 System-on-Chip aboard a Crazyflie nano-UAV. Then, to cope with in-field unexpected obstacles, we leverage a global+local path planner based on the A* algorithm. The global path planner determines the best route for the nano-UAV to sweep the entire area, while the local one runs up to 50 Hz aboard our nano-UAV and prevents collision by adjusting the short-distance path. Finally, we demonstrate with in-simulator experiments that once a 25 nano-UAVs fleet has combed a 200 {\texttimes} 200 m vineyard, collected information can be used to plan the best path for the tractor, visiting all and only required hotspots. In this scenario, our efficient transportation system, compared to a traditional single-ground vehicle performing both inspection and treatment, can save up to 20 h working time.

On the Regularization of Learnable Embeddings for Time Series Forecasting

Luca Butera, Giovanni De Felice, Andrea Cini, Cesare Alippi

Transactions on Machine Learning Research · 2025

In forecasting multiple time series, accounting for the individual features of each sequence can be challenging. To address this, modern deep learning methods for time series analysis combine a shared (global) model with local layers, specific to each time series, often implemented as learnable embeddings. Ideally, these local embeddings should encode meaningful representations of the unique dynamics of each sequence. However, when these are learned end-to-end as parameters of a forecasting model, they may end up acting as mere sequence identifiers. Shared processing blocks may then become reliant on such identifiers, limiting their transferability to new contexts. In this paper, we address this issue by investigating methods to regularize the learning of local learnable embeddings for time series processing. Specifically, we perform the first extensive empirical study on the subject and show how such regularizations consistently improve performance in widely adopted architectures. Furthermore, we show that methods attempting to prevent the co-adaptation of local and global parameters by means of embeddings perturbation are particularly effective in this context. In this regard, we include in the comparison several perturbation-based regularization methods, going as far as periodically resetting the embeddings during training. The obtained results provide an important contribution to understanding the interplay between learnable local parameters and shared processing layers: a key challenge in modern time series processing models and a step toward developing effective foundation models for time series.

2024

Object-Centric Relational Representations for Image Generation

Luca Butera, Andrea Cini, Alberto Ferrante, Cesare Alippi

Transactions on Machine Learning Research · 2024

Conditioning image generation on specific features of the desired output is a key ingredient of modern generative models. However, existing approaches lack a general and unified way of representing structural and semantic conditioning at diverse granularity levels. This paper explores a novel method to condition image generation, based on object-centric relational representations. In particular, we propose a methodology to condition the generation of objects in an image on the attributed graph representing their structure and the associated semantic information. We show that such architectural biases entail properties that facilitate the manipulation and conditioning of the generative process and allow for regularizing the training procedure. The proposed conditioning framework is implemented by means of a neural network that learns to generate a 2D, multi-channel, layout mask of the objects, which can be used as a soft inductive bias in the downstream generative task. To do so, we leverage both 2D and graph convolutional operators. We also propose a novel benchmark for image generation consisting of a synthetic dataset of images paired with their relational representation. Empirical results show that the proposed approach compares favorably against relevant baselines.

A Deep Learning-based Pest Insect Monitoring System for Ultra-low Power Pocket-sized Drones

Luca Crupi, Luca Butera, Alberto Ferrante, Daniele Palossi

2024 20th International Conference on Distributed Computing in Smart Systems and the Internet of Things (DCOSS-IoT) · 2024

Smart farming and precision agriculture represent game-changer technologies for efficient and sustainable agribusiness. Miniaturized palm-sized drones can act as flexible smart sensors inspecting crops, looking for early signs of potential pest outbreaking. However, achieving such an ambitious goal requires hardware-software codesign to develop accurate deep learning (DL) detection models while keeping memory and computational needs under an ultra-tight budget, i.e., a few MB on-chip memory and a few 100s mW power envelope. This work presents a novel vertically integrated solution featuring two ultra-low power System-on-Chips (SoCs), i.e., the dual-core STM32H74 and a multi-core GWT GAP9, running two State-of-the-Art DL models for detecting the Popillia japonica bug. We fine-tune both models for our image-based detection task, quantize them in 8-bit integers, and deploy them on the two SoCs. On the STM32H74, we deploy a FOMO-MobileNetV2 model, achieving a mean average precision (mAP) of 0.66 and running at 16.1 frame/s within 498 mW. While on the GAP9 SoC, we deploy a more complex SSDLite-MobileNetV3, which scores an mAP of 0.79 and peaks at 6.8 frame/s within 33 mW. Compared to a top-notch RetinaNet-ResNet101-FPN full-precision baseline, which requires 14.9× more memory and 300× more operations per inference, our best model drops only 15% in mAP, paving the way toward autonomous palm-sized drones capable of lightweight and precise pest detection.

2022

Precise Agriculture: Effective Deep Learning Strategies to Detect Pest Insects

Luca Butera, Alberto Ferrante, Mauro Jermini, Mauro Prevostini, Cesare Alippi

IEEE/CAA Journal of Automatica Sinica · 2022

Pest insect monitoring and control is crucial to ensure a safe and profitable crop growth in all plantation types, as well as guarantee food quality and limited use of pesticides. We aim at extending traditional monitoring by means of traps, by involving the general public in reporting the presence of insects by using smartphones. This includes the largely unexplored problem of detecting insects in images that are taken in noncontrolled conditions. Furthermore, pest insects are, in many cases, extremely similar to other species that are harmless. Therefore, computer vision algorithms must not be fooled by these similar insects, not to raise unmotivated alarms. In this work, we study the capabilities of state-of-the-art (SoA) object detection models based on convolutional neural networks (CNN) for the task of detecting beetle-like pest insects on nonhomogeneous images taken outdoors by different sources. Moreover, we focus on disambiguating a pest insect from similar harmless species. We consider not only detection performance of different models, but also required computational resources. This study aims at providing a baseline model for this kind of tasks. Our results show the suitability of current SoA models for this application, highlighting how FasterRCNN with a MobileNetV3 backbone is a particularly good starting point for accuracy and inference execution latency. This combination provided a mean average precision score of 92.66% that can be considered qualitatively at least as good as the score obtained by other authors that adopted more specific models.

2021

The human nasal cavity: towards the optimal surgery with CFD and Machine Learning

Andrea Schillaci, Luca Butera, Gianluca Romani, Carlotta Pipolo, Giovanni Felisati, Marcello Restelli, Giacomo Boracchi, Maurizio Quadrio

25th International Congress of Theoretical and Applied Mechanics (XXV ICTAM) · 2021

Nasal breathing difficulties are a common condition, and their treatment often requires surgery. Unfortunately, procedures are designed and carried out mostly based on the surgeon’s experience; data supporting surgical choices are lacking. Computational Fluid Dynamics (CFD), by naturally accessesing functional properties of the human nose, improves the understanding of its flow field. However, a detailed flow prediction alone does not immediately lead to identifying the best surgical maneuver. We intend to leverage Machine Learning (ML) to bridge this gap and infer functional information from CFD data. The present study is preliminary and uses rather crude anatomical and computational models; however, its results demonstrate the potential of a ML model trained on CFD data.

2019

Nasal pathology assessment through supervised learning on computational fluid dynamics data : a preliminary study

Luca Butera

Politecnico di Milano · 2019

Otolaryngology, in particular rhinology, is a field of medicine in which diagnosis tools and techniques are mostly bound to the analysis of nasal cavities' morphology and overall airflow quality through outdated technologies, if compared to other medical branches. However, in recent years, with the growth and spread of computational power at hand, CFD (Computational Fluid Dynamics), has shown the possibility to enhance the rhinologic diagnosis with highly detailed simulations of patients' nasal cavities airflow. This powerful tool, however, poses the challenge of results interpretation, given the high dimensionality of data produced by these simulations and the lack of understanding, by both medical professionals and engineers, of the non-trivial correlation between these data and the presence of a pathological condition. In this setting ML (Machine Learning) could represent the key to craft a diagnosis tool which produces results of easier interpretability for the doctor. The aim of this thesis is to investigate the presence of information in CFD data that can relate to the existence of a pathological condition. To do so we created a simplified parametric nose model, which we modified accordingly to both pathological and non-pathological traits, and executed fluid dynamics simulations over it. From simulation data we extracted various features to finally propose the study of their predictivity in a Supervised Learning setting.