[AAAI 2023 (Oral)] CrissCross: Self-Supervised Audio-Visual Representation Learning with Relaxed Cross-Modal Synchronicity
-
Updated
Jul 11, 2023 - Python
[AAAI 2023 (Oral)] CrissCross: Self-Supervised Audio-Visual Representation Learning with Relaxed Cross-Modal Synchronicity
End-to-end pipeline for training a custom keyword detection model with TensorFlow & TFLite expor
Simplified PyTorch implementation of audio classification, support multi-gpu training and validating, automatic mixed precision training, knowledge distillation etc.
Model Deployment for HEAR4U Bangkit Capstone Project
A novel, model-agnostic Explainable AI (XAI) framework for audio classification. Introduces RISE-SPEC, RISE-WAVE, and RISE-AUDIO, demonstrating superior interpretability over baseline methods like RISE, LIME, and Grad-CAM.
Tuned SVM for classifying 10 classes form ESC-50 dataset using acoustic features manually extracted. The final model achieved 86 percent accuracy.
AI-powered classroom audio event recognition using YAMNet embeddings and XGBoost achieving 91.25% accuracy on the ESC-50 dataset.
SoundGuard is a GenAI agent that detects emergency sounds, explains what it hears, and responds like a smart assistant — built with YAMNet, Gradio, Google Cloud, and deployed on Hugging Face Spaces.
This project focuses on the ESC50 Challenge. The ESC-50 dataset is a labeled collection of 2000 environmental audio recordings suitable for benchmarking methods of environmental sound classification.
This repository contains the implementation of Environmental Sound Classification on the ESC-50 dataset using the ACDNet.
To associate your repository with the esc50 topic, visit your repo's landing page and select "manage topics."