Abstract
Trajectory prediction is essential for autonomous driving to anticipate traffic participants’ movements. While recent methods incorporate road topology, they often rely on heavy architectures, limiting real-time deployment. This study proposes a lightweight adversarial framework for multimodal trajectory prediction, integrating spatio-temporal encoding of agents and lanes, hierarchical attention-based scene fusion, a Bézier-parameterized generator, and a context-aware discriminator. This approach captures agent-road interactions and ensures trajectory realism with only 1.9M parameters. Experimental results demonstrate the superior speed and accuracy of the proposed method compared to existing approaches for trajectory prediction. The proposed method significantly enhances prediction accuracy, achieving results of minADE6 0.76 and minFDE6 1.12 on the Argoverse 1 dataset. In comparison to the state-of-the-art baseline model, there are notable improvements in minADE6 and minFDE6 by 0.01 and 0.03, respectively. Furthermore, our method generalizes well to the Argoverse 2 dataset, showing strong performance in more diverse and challenging environments.
Keywords
Get full access to this article
View all access options for this article.
