TopGit / yandexdataschool/speech_course SNAPSHOT GITHUB
yandexdataschool/speech_course Y TopGit repo profile for yandexdataschool/speech_course, with GitHub repository stats and README context. Đánh giá Readme Prompt Clone Nhắc đến Bình luận
ĐÁNH GIÁ NHANH
yandexdataschool/speech_course — dự án mã nguồn mở — đang có 345 sao GitHub. YSDA course in Speech Processing.
Tóm tắt dựng từ metadata GitHub của chính dự án — chưa có bài review TopGit. Trang sẽ tự động cập nhật khi bài review đầy đủ được xuất bản.
VÌ SAO CHƯA CÓ REVIEW
TopGit viết bài đầy đủ cho repo có nhiều sao nhất và được yêu cầu nhiều nhất. Trang này là snapshot trong thời gian chờ — xem README gốc ở tab READ ME.
Snapshot Stars ★ 345
Forks ⑂ 104
Language Jupyter Notebook
Topic —
License MIT
Homepage —
Xem cách chúng tôi review ↗
Cộng tác viên hàng đầu Xem cộng tác viên hàng đầu YSDA Speech Processing Course
Materials for each week are in ./week* folders
Course program
Week 1: Slides | Lecture | Seminar
Lecture: Intro to Digital Signal Processing (DSP)
Seminar: Implement DSP pipeline
Homework (5pt): Implement mel-spectrogram transformations
Week 2: Slides | Lecture | Seminar
Lecture: Introduction to speech NN discriminative models. Voice Activity Detection (VAD) and Sound Event Detection (SED) tasks
Seminar: Train VAD models, intro to the homework
Homework (15pt): Train SED models; (3pt bonus) SED models vibecoding
Week 3: Slides | Lecture | Seminar
Lecture: Keyword Spotting and Speech Biometrics tasks
Seminar: Train Biometrics model and look at embeddings
Homework (20pt): Train Biometrics model ECAPA-TDNN with contrastive loss
Week 4 Slides | Lecture+Seminar
Lecture: Speech Recognition I
Seminar: CTC forward-backward, soft alignment
Homework (10pt): CTC/RNN-T decoding, RNN-T forward-backward
Week 5 Slides | Lecture | Seminar
Lecture: Pretraining in Speech Recognition
Seminar: Speech Pretraining - quantization and losses
Homework (5pt): Speech Pretraining
Week 6 Slides | Lecture
Leсture: ASR Inference
Homework (5pt): Implement streaming inference
Week 7 Slides | Lecture
Lecture: Intro to TTS. Normalisation, Tasks, Metrics
Week 8 Slides | Lecture
Lecture: Tacotron2, FastPitch, HiFiGAN
Seminar (5pt): Implement pitch estimation
Homework (10pt): Implement FastPitch
Week 9 Slides | Lecture | Seminar
Lecture: Quantisation and Neural codecs
Seminar: Implement several quantisation methods
Homework (10pt): Implement more advanced quantisations and audio codecs
Week 10 Slides | Lecture
Lecture: Diffusions and transformers for voice cloning
Homework (10pt): Implement slow-fast transformer inference
Week 11 Slides | Lecture
Lecture: Spoken dialogue models
Week 13 Slides | Lecture
Lecture: AEC and beamforming
Homework (5pt): Implement AEC
Course program for spring 2025
Week 1: Slides | Lecture | Seminar
Lecture: Intro to Digital Signal Processing (DSP)
Seminar: Implement DSP pipeline
Homework (5pt): Implement mel-spectrogram transformations
Week 2 Slides | Lecture | Seminar:
Lecture: Introduction to speech NN discriminative models. Voice Activity Detection (VAD) and Sound Event Detection (SED) tasks
Seminar: Train VAD models
Homework (15pt): Train SED models
Week 3 Slides | Lecture | Seminar:
Lecture: Keyword Spotting and Speech Biometrics tasks
Seminar: Train Biometrics model and look at embeddings
Homework (20pt): Train Biometrics model ECAPA-TDNN with contrastive loss
Week 4 Slides | Lecture | Seminar:
Lecture: Speech Recognition I
Seminar: CTC forward-backward, soft alignment
Homework (10pt): CTC/RNN-T decoding, RNN-T forward-backward
Week 5 Slides | Lecture | Seminar:
Lecture: Speech Recognition II, Pretraining
Homework (5pt): Finetune Wav2Vec2
Week 6: Slides | Lecture | Seminar
Lecture: ASR Inference
Seminar: Streaming ASR
Homework (5pt): Seminar continuation
Week 7: Slides | Lecture
Lecture: Text-to-Speech I, intro, preprocessor, metrics
Week 8: Slides | Lecture
Lecture: Text-to-Speech II, Acoustic models and vocoding
Seminar (5pt): Pitch estimation, Monotonic Alignment Search for phoneme duration estimation
Homework (10pt): Train FastPitch model
Week 9: Slides | Lecture | Seminar
Lecture: Text-to-Speech III, Codecs
Seminar: Vector Quantizaton, Residual Vector Quantization
Week 10: Slides | Lecture
Lecture: Text-to-Speech IV, Tortoise and other tranformers for TTS
Homework (15pt): write inference for CLM with two transformers
Week 11: Slides | Lecture
Lecture: Multimodality, How to build a big GPT with voice capabilities
Week 12: Slides | Lecture | Seminar
Lecture: noise reduction
Seminar: Streaming STFT and ISTFT
Homework (15pt): Noise reduction model implementation
Week 13: Slides 1 | Slides 2 | Lecture+Seminar
Lecture: Acoustic Echo Cancelation (AEC) and Beamforming
Homework (5pt): Basic AEC implementation
Course program for spring 2024
Week 1: Slides | Lecture | Seminar
Lecture: Intro to Digital Signal Processing (DSP)
Seminar: Implement DSP pipeline
Week 2: Slides | Lecture | Seminar
Lecture: Introduction to speech NN discriminative models. Voice Activity Detection (VAD) and Sound Event Detection (SED) tasks
Seminar: Train VAD models
Homework: Train SED models
Week 3: Slides | Lecture | Seminar
Lecture: Keyword Spotting and Speech Biometrics tasks
Seminar: Train Biometrics model and look at embeddings
Homework: Train Biometrics model to better quality
Week 4: Slides | Lecture | Seminar
Lecture: Speech Recognition I
Seminar: Metrics and augmentations for speech recognition
Homework: Implement CTC algorithm
Week 5: Slides | Lecture
Lecture: Speech Recognition II, Pretraining
Homework: Finetune Wav2Vec2
Week 6: Slides | Lecture
Lecture: Text-to-Speech I, intro, preprocessor, metrics
Week 7: Slides | Lecture
Lecture: Text-to-Speech II, Acoustic models
Seminar: Pitch estimation, Monotonic Alignment Search for phoneme duration estimation
Homework: Train FastPitch model
Week 8: Slides, p1 | Lecture, p1 | Slides, p2 | Lecture, p2 | Seminar
Lecture, p1: Text-to-Speech III, Vocoding
Lecture, p2: Vector Quantization, Codecs
Seminar: Vector Quantizaton, Residual Vector Quantization
Week 9: Slides | Lecture, p1 | Lecture, p2
Lecture: Tranformers for TTS
Homework: write inference for pre-trained transformer
Week 10: Slides | Lecture | Seminar
Lecture: noise reduction
Seminar: Streaming STFT and ISTFT
Homework: Noise reduction model implementation
Week 11: Slides | Lecture
Lecture: Acoustic Echo Cancelation (AEC) and Beamforming
Week 12: Slides | Lecture | Seminar
Lecture: ASR Inference
Seminar: Streaming ASR
Week 13: Slides | Lecture
Lecture: Flow based TTS + Voice Conversion
Contributors & course staff
Current:
Pavel Mazaev - VAD, SED
Aman Syayfetdinov - spotter, biometry
Daniil Volgin - ASR
Dzmitry Soupel - ASR
Stepan Kargaltsev - ASR
Roma Kail - TTS
Arina Shchegortsova TTS
Vladimir Gogoryan - TTS
Ravil Khisamov - VQE
Anton Porfirev - AEC
Previous iteration:
Evgeniia Elistratova - TTS lectures, seminars and homeworks
Vladimir Platonov - TTS lectures
Andrey Malinin - Course admin, lectures, seminars, homeworks
Vladimir Kirichenko - lectures, seminars, homeworks
Segey Dukanov - lecures, seminars, homeworks
Evgenii Shabalin - lecture and homework on conversion
Mikhail Andreev - ASR
Alex Rak - VAD, SED, spotter, biometry
CÂU HỎI THƯỜNG GẶP
Trả lời nhanh yandexdataschool/speech_course có bao nhiêu sao? yandexdataschool/speech_course có 345 sao GitHub — tải lại trang để xem số mới nhất, hoặc xem trực tiếp github.com/yandexdataschool/speech_course. TopGit phản chiếu số sao của GitHub nhưng không cam kết đến từng phút.
yandexdataschool/speech_course có những chủ đề gì? GitHub topics của yandexdataschool/speech_course: "asr", "dsp", "tts", "vqe". TopGit xếp repo vào nhóm mã nguồn mở.
yandexdataschool/speech_course còn đang phát triển không? Commit gần nhất trên yandexdataschool/speech_course là 3 tháng trước (theo timestamp GitHub). Repo có 104 fork — một chỉ báo về mức độ quan tâm của cộng đồng.
yandexdataschool/speech_course là gì? yandexdataschool/speech_course (yandexdataschool/speech_course) là dự án Jupyter Notebook trên GitHub. Theo mô tả gốc: YSDA course in Speech Processing.
yandexdataschool/speech_course viết bằng ngôn ngữ gì? yandexdataschool/speech_course chủ yếu viết bằng Jupyter Notebook. Trường "language" của GitHub dựa trên phần lớn byte ở nhánh mặc định.
Đọc đầy đủ README ở tab phía trên.
speech_course có đáng để bạn bỏ thời gian? ChatGPT, Claude và Perplexity đều đọc được trang này. Hỏi thử xem họ nghĩ gì về speech_course.