| 1 |
Chen, S., Wu, Y., Wang, C., Liu, S., Tompkins, D., Chen, Z., and Wei, F. (2022). Beats: Audio pre-training with acoustic tokenizers. arXiv preprint arXiv:2212.09058.
|
|
| 2 |
Gemmeke, J. F., Ellis, D. P. W., Freedman, D., Jansen, A., Lawrence, W., Moore, R. C., Plakal, M., and Ritter, M. (2017). AudioSet: An ontology and human-labeled dataset for audio events. In Proc. IEEE ICASSP, pages 776–780.
|
|
| 3 |
Gong, Y., Chung, Y.-A., and Glass, J. (2021). AST: Audio spectrogram transformer. In Proc. Interspeech, pages 571–575.
|
|
| 4 |
Gu, A. and Dao, T. (2023). Mamba: Linear-time sequence modeling with selective state spaces. arXiv:2312.00752.
|
|
| 5 |
Kong, Q., Cao, Y., Iqbal, T., Wang, Y., Wang, W., and Plumbley, M. D. (2020). PANNs: Large-scale pretrained audio neural networks for audio pattern recognition. IEEE/ACM Trans. Audio, Speech, Lang. Process., 28:2880–2894.
|
|
| 6 |
Mohmmad, S. and Sanampudi, S. K. (2024). Exploring current research trends in sound event detection: a systematic literature review. Multimedia Tools and Applications, 83:84699–84741.
|
|
| 7 |
Moreira Souza, A., Moreira, G. A., and Pulcinelli, L. E. G. (2025). A Comparative Analysis of Denoising Methods for Deep Learning-Based Audio Event Detection in Noisy Agricultural Environments. In Anais do XL Simpósio Brasileiro de Banco de Dados (SBBD 2025), pages 942–948, Brasil. Sociedade Brasileira de Computação SBC.
|
|
| 8 |
Mu, D., Zhang, Z., and Yue, H. (2024). MFF-EINV2: Multi-scale feature fusion across spectral-spatial-temporal domains for sound event localization and detection. In Proc. Interspeech, pages 92–96.
|
|
| 9 |
Schmid, F., Koutini, K., and Widmer, G. (2023). Efficient large-scale audio tagging via transformer-to-CNN knowledge distillation. In Proc. IEEE ICASSP, pages 1–5.
|
|
| 10 |
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2015). Going deeper with convolutions. In Proc. IEEE CVPR, pages 1–9.
|
|
| 11 |
Turab, M., Kumar, T., Bendechache, M., and Saber, T. (2022). Investigating multi-feature selection and ensembling for audio classification. arXiv:2206.07511.
|
|
| 12 |
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Information Processing Systems, volume 30, pages 5998–6008.
|
|
| 13 |
Yadav, S. and Tan, Z.-H. (2024). Audio mamba: Selective state spaces for self-supervised audio representations. In Proc. Interspeech, pages 552–556.
|
|
| 14 |
Zhang, Y., Huang, D., and Togneri, R. (2025). Pseudo strong labels from frame-level predictions for weakly supervised sound event detection. arXiv:2501.03740.
|
|