Kyubyong/nlp_tasks
Kyubyong/nlp_tasks is one of the AI-powered repositories TopGit tracks, currently at 3.0k stars. Natural Language Processing Tasks and References
Snapshot summary built from the project's own GitHub metadata — there's no written TopGit review yet. The page will update automatically when a full review is published.
TopGit writes full reviews for the most-starred, most-requested repositories. This page is a snapshot until then — see the READ ME tab for the original README in full.
Snapshot
Top contributors
Show top contributors
Natural Language Processing Tasks and Selected References
I've been working on several natural language processing tasks for a long time. One day, I felt like drawing a map of the NLP field where I earn a living. I'm sure I'm not the only person who wants to see at a glance which tasks are in NLP.
I did my best to cover as many as possible tasks in NLP, but admittedly this is far from exhaustive purely due to my lack of knowledge. And selected references are biased towards recent deep learning accomplishments. I expect these serve as a starting point when you're about to dig into the task. I'll keep updating this repo myself, but what I really hope is you collaborate on this work. Don't hesitate to send me a pull request!
Oct. 13, 2017.
by Kyubyong
Reviewed and updated by YJ Choe on Oct. 18, 2017.
Anaphora Resolution
- See Coreference Resolution
Automated Essay Scoring
PAPERAutomatic Text Scoring Using Neural NetworksPAPERA Neural Approach to Automated Essay ScoringCHALLENGEKaggle: The Hewlett Foundation: Automated Essay ScoringPROJECTEASE (Enhanced AI Scoring Engine)
Automatic Speech Recognition
WIKISpeech recognitionPAPERDeep Speech 2: End-to-End Speech Recognition in English and MandarinPAPERWaveNet: A Generative Model for Raw AudioPROJECTA TensorFlow implementation of Baidu's DeepSpeech architecturePROJECTSpeech-to-Text-WaveNet : End-to-end sentence level English speech recognition using DeepMind's WaveNetCHALLENGEThe 5th CHiME Speech Separation and Recognition ChallengeDATAThe 5th CHiME Speech Separation and Recognition ChallengeDATACSTR VCTK CorpusDATALibriSpeech ASR corpusDATASwitchboard-1 Telephone Speech CorpusDATATED-LIUM CorpusDATAOpen Speech and Language ResourcesDATACommon Voice
Automatic Summarisation
WIKIAutomatic summarizationBOOKAutomatic Text SummarizationPAPERText Summarization Using Neural NetworksPAPERRanking with Recursive Neural Networks and Its Application to Multi-Document SummarizationDATAText Analytics Conferences (TAC)DATADocument Understanding Conferences (DUC)
Coreference Resolution
INFOCoreference ResolutionPAPERDeep Reinforcement Learning for Mention-Ranking Coreference ModelsPAPERImproving Coreference Resolution by Learning Entity-Level Distributed RepresentationsCHALLENGECoNLL 2012 Shared Task: Modeling Multilingual Unrestricted Coreference in OntoNotesCHALLENGECoNLL 2011 Shared Task: Modeling Unrestricted Coreference in OntoNotesCHALLENGESemEval 2018 Task 4: Character Identification on Multiparty Dialogues
Entity Linking
- See Named Entity Disambiguation
Grammatical Error Correction
PAPERA Multilayer Convolutional Encoder-Decoder Neural Network for Grammatical Error CorrectionPAPERNeural Network Translation Models for Grammatical Error CorrectionPAPERAdapting Sequence Models for Sentence CorrectionCHALLENGECoNLL-2013 Shared Task: Grammatical Error CorrectionCHALLENGECoNLL-2014 Shared Task: Grammatical Error CorrectionDATANUS Non-commercial research/trial corpus licenseDATALang-8 Learner CorporaDATACornell Movie--Dialogs CorpusPROJECTDeep Text CorrectorPRODUCTdeep grammar
Grapheme To Phoneme Conversion
PAPERGrapheme-to-Phoneme Models for (Almost) Any LanguagePAPERPolyglot Neural Language Models: A Case Study in Cross-Lingual Phonetic Representation LearningPAPERMultitask Sequence-to-Sequence Models for Grapheme-to-Phoneme ConversionPROJECTSequence-to-Sequence G2P toolkitPROJECTg2p_en: A Simple Python Module for English Grapheme To Phoneme ConversionDATAMultilingual Pronunciation Data
Humor and Sarcasm Detection
PAPERAutomatic Sarcasm Detection: A SurveyPAPERMagnets for Sarcasm: Making Sarcasm Detection Timely, Contextual and Very PersonalPAPERSarcasm Detection on Twitter: A Behavioral Modeling ApproachCHALLENGESemEval-2017 Task 6: #HashtagWars: Learning a Sense of HumorCHALLENGESemEval-2017 Task 7: Detection and Interpretation of English PunsDATASarcastic comments from RedditDATASarcasm Corpus V2DATASarcasm Amazon Reviews Corpus
Language Grounding
WIKISymbol grounding problemPAPERThe Symbol Grounding ProblemPAPERFrom phonemes to images: levels of representation in a recurrent neural model of visually-grounded language learningPAPEREncoding of phonology in a recurrent neural model of grounded speechPAPERGated-Attention Architectures for Task-Oriented Language GroundingPAPERSound-Word2Vec: Learning Word Representations Grounded in SoundsCOURSELanguage Grounding to Vision and ControlWORKSHOPLanguage Grounding for Robotics
Language Guessing
- See Language Identification
Language Identification
WIKILanguage identificationPAPERAUTOMATIC LANGUAGE IDENTIFICATION USING DEEP NEURAL NETWORKSPAPERNatural Language Processing with Small Feed-Forward NetworksCHALLENGE2015 Language Recognition Evaluation
Language Modeling
WIKILanguage modelTOOLKITKenLM Language Model ToolkitPAPERDistributed Representations of Words and Phrases and their CompositionalityPAPERGenerating Sequences with Recurrent Neural NetworksPAPERCharacter-Aware Neural Language ModelsTHESISStatistical Language Models Based on Neural NetworksDATAPenn TreebankTUTORIALTensorFlow Tutorial on Language Modeling with Recurrent Neural Networks
Language Recognition
- See Language Identification
Lemmatisation
WIKILemmatisationPAPERJoint Lemmatization and Morphological Tagging with LEMMINGTOOLKITWordNet LemmatizerDATATreebank-3
Lip-reading
WIKILip readingPAPERLipNet: End-to-End Sentence-level LipreadingPAPERLip Reading Sentences in the WildPAPERLarge-Scale Visual Speech RecognitionPROJECTLip Reading - Cross Audio-Visual Recognition using 3D Convolutional Neural NetworksPRODUCTLiopaDATAThe GRID audiovisual sentence corpusDATAThe BBC-Oxford 'Multi-View Lip Reading Sentences' (MV-LRS) Dataset
Machine Translation
PAPERNeural Machine Translation by Jointly Learning to Align and TranslatePAPERNeural Machine Translation in Linear TimePAPERAttention Is All You NeedPAPERSix Challenges for Neural Machine TranslationPAPERPhrase-Based & Neural Unsupervised Machine TranslationCHALLENGEACL 2014 NINTH WORKSHOP ON STATISTICAL MACHINE TRANSLATIONCHALLENGEEMNLP 2017 SECOND CONFERENCE ON MACHINE TRANSLATION (WMT17)DATAOpenSubtitles2016DATAWIT3: Web Inventory of Transcribed and Translated TalksDATAThe QCRI Educational Domain (QED) CorpusPAPERMulti-task Sequence to Sequence LearningPAPERUnsupervised Pretraining for Sequence to Sequence LearningPAPERGoogle’s Multilingual Neural Machine Translation System: Enabling Zero-Shot TranslationTOOLKITSubword Neural Machine Translation with Byte Pair Encoding (BPE)TOOLKITMulti-Way Neural Machine TranslationTOOLKITOpenNMT: Open-Source Toolkit for Neural Machine Translation
Morphological Inflection Generation
WIKIInflectionPAPERMorphological Inflection Generation Using Character Sequence to Sequence LearningCHALLENGESIGMORPHON 2016 Shared Task: Morphological ReinflectionDATAsigmorphon2016
Named Entity Disambiguation
WIKIEntity linkingPAPERRobust and Collective Entity Disambiguation through Semantic Embeddings
Named Entity Recognition
WIKINamed-entity recognitionPAPERNeural Architectures for Named Entity RecognitionPROJECTOSU Twitter NLP ToolsCHALLENGENamed Entity Recognition in TwitterCHALLENGECoNLL 2002 Language-Independent Named Entity RecognitionCHALLENGEIntroduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity RecognitionDATACoNLL-2002 NER corpusDATACoNLL-2003 NER corpusDATANUT Named Entity Recognition in Twitter Shared taskTOOLKITStanford Named Entity Recognizer
Paraphrase Detection
PAPERDynamic Pooling and Unfolding Recursive Autoencoders for Paraphrase DetectionPROJECTParalex: Paraphrase-Driven Learning for Open Question AnsweringCHALLENGESemEval-2015 Task 1: Paraphrase and Semantic Similarity in TwitterDATAMicrosoft Research Paraphrase CorpusDATAMicrosoft Research Video Description CorpusDATAPascal DatasetDATAFlickr DatasetDATAThe SICK data setDATAPPDB: The Paraphrase DatabaseDATAWikiAnswers Paraphrase Corpus
Paraphrase Generation
PAPERNeural Paraphrase Generation with Stacked Residual LSTM NetworksDATANeural Paraphrase Generation with Stacked Residual LSTM NetworksCODENeural Paraphrase Generation with Stacked Residual LSTM NetworksPAPERA Deep Generative Framework for Paraphrase GenerationPAPERParaphrasing Revisited with Neural Machine Translation
Parsing
WIKIParsingTOOLKITThe Stanford Parser: A statistical parserTOOLKITspaCy parserPAPERGrammar as a Foreign LanguagePAPERA fast and accurate dependency parser using neural networksPAPERUniversal Semantic ParsingCHALLENGECoNLL 2017 Shared Task: Multilingual Parsing from Raw Text to Universal DependenciesCHALLENGECoNLL 2016 Shared Task: Multilingual Shallow Discourse ParsingCHALLENGECoNLL 2015 Shared Task: Shallow Discourse ParsingCHALLENGESemEval-2016 Task 8: The meaning representations may be abstract, but this task is concrete!
Part-of-speech Tagging
WIKIPart-of-speech taggingPAPERMultilingual Part-of-Speech Tagging with Bidirectional Long Short-Term Memory Models and Auxiliary LossPAPERUnsupervised Part-Of-Speech Tagging with Anchor Hidden Markov ModelsDATATreebank-3TOOLKITnltk.tag package
Pinyin-To-Chinese Conversion
WIKIPinyin input methodPAPERNeural Network Language Model for Chinese Pinyin Input Method EnginePROJECTNeural Chinese Transliterator
Question Answering
WIKIQuestion answeringPAPERAsk Me Anything: Dynamic Memory Networks for Natural Language ProcessingPAPERDynamic Memory Networks for Visual and Textual Question AnsweringCHALLENGETREC Question Answering TaskCHALLENGENTCIR-8: Advanced Cross-lingual Information Access (ACLIA)CHALLENGECLEF Question Answering TrackCHALLENGESemEval-2017 Task 3: Community Question AnsweringCHALLENGESemEval-2018 Task 11: Machine Comprehension using Commonsense KnowledgeDATAMS MARCO: Microsoft MAchine Reading COmprehension DatasetDATAMaluuba NewsQADATASQuAD: 100,000+ Questions for Machine Comprehension of TextDATAGraphQuestions: A Characteristic-rich Question Answering DatasetDATAStory Cloze Test and ROCStories CorporaDATAMicrosoft Research WikiQA CorpusDATADeepMind Q&A DatasetDATAQASentDATATextbook Question Answering
Relationship Extraction
WIKIRelationship extractionPAPERA deep learning approach for relationship extraction from interaction context in social manufacturing paradigmCHALLENGESemEval-2018 task 7 Semantic Relation Extraction and Classification in Scientific Papers
Semantic Role Labeling
WIKISemantic role labelingBOOKSemantic Role LabelingPAPEREnd-to-end Learning of Semantic Role Labeling Using Recurrent Neural NetworksPAPERNeural Semantic Role Labeling with Dependency Path EmbeddingsPAPERDeep Semantic Role Labeling: What Works and What's NextCHALLENGECoNLL-2005 Shared Task: Semantic Role LabelingCHALLENGECoNLL-2004 Shared Task: Semantic Role LabelingTOOLKITIllinois Semantic Role Labeler (SRL)DATACoNLL-2005 Shared Task: Semantic Role Labeling
Sentence Boundary Disambiguation
WIKISentence boundary disambiguationPAPERA Quantitative and Qualitative Evaluation of Sentence Boundary Detection for the Clinical DomainTOOLKITNLTK TokenizersDATAThe British National CorpusDATASwitchboard-1 Telephone Speech Corpus
Sentiment Analysis
WIKISentiment analysisINFOAwesome Sentiment AnalysisCHALLENGEKaggle: UMICH SI650 - Sentiment ClassificationCHALLENGESemEval-2017 Task 4: Sentiment Analysis in TwitterCHALLENGESemEval-2017 Task 5: Fine-Grained Sentiment Analysis on Financial Microblogs and NewsPROJECTSenticNetPROJECTStanford NLP Group Sentiment AnalysisDATAMulti-Domain Sentiment Dataset (version 2.0)DATAStanford Sentiment TreebankDATATwitter Sentiment CorpusDATATwitter Sentiment Analysis Training CorpusDATAAFINN: List of English words rated for valence
Sign Language Recognition/Translation
PAPERVideo-based Sign Language Recognition without Temporal SegmentationPAPERSubUNets: End-to-end Hand Shape and Continuous Sign Language RecognitionDATARWTH-PHOENIX-WeatherDATAASLLRPPROJECTSignAll
Singing Voice Synthesis
PAPERSinging voice synthesis based on deep neural networksPAPERA Neural Parametric Singing Synthesizer Modeling Timbre and Expression from Natural SongsPRODUCTVOCALOID: voice synthesis technology and software developed by YamahaCHALLENGESpecial Session Interspeech 2016 Singing synthesis challenge "Fill-in the Gap"
Social Science Applications
WORKSHOPNLP+CSS: Workshops on Natural Language Processing and Computational Social ScienceTOOLKITMen Also Like Shopping: Reducing Gender Bias Amplification using Corpus-level ConstraintsTOOLKITOnline Variational Bayes for Latent Dirichlet Allocation (LDA)GROUPThe University of Chicago Knowledge Lab
Source Separation
WIKISource separationPAPERFrom Blind to Guided Audio Source SeparationPAPERJoint Optimization of Masks and Deep Recurrent Neural Networks for Monaural Source SeparationCHALLENGESignal Separation Evaluation Campaign (SiSEC)CHALLENGECHiME Speech Separation and Recognition Challenge
Speaker Authentication
- See Speaker Verification
Speaker Diarisation
WIKISpeaker diarisationPAPERDNN-based speaker clustering for speaker diarisationPAPERUnsupervised Methods for Speaker Diarization: An Integrated and Iterative ApproachPAPERAudio-Visual Speaker Diarization Based on Spatiotemporal Bayesian FusionCHALLENGERich Transcription Evaluation
Speaker Recognition
WIKISpeaker recognitionPAPERA NOVEL SCHEME FOR SPEAKER RECOGNITION USING A PHONETICALLY-AWARE DEEP NEURAL NETWORKPAPERDEEP NEURAL NETWORKS FOR SMALL FOOTPRINT TEXT-DEPENDENT SPEAKER VERIFICATIONPAPERDeep Speaker: an End-to-End Neural Speaker Embedding SystemPROJECTVoice Vector: which of the Hollywood stars is most similar to my voice?CHALLENGENIST Speaker Recognition Evaluation (SRE)INFOAre there any suggestions for free databases for speaker recognition?DATAVoxCeleb2: Deep Speaker Recognition
Speech Reading
- See Lip-reading
Speech Recognition
- See Automatic Speech Recognition
Speech Segmentation
WIKISpeech_segmentationPAPERWord Segmentation by 8-Month-Olds: When Speech Cues Count More Than StatisticsPAPERUnsupervised Word Segmentation and Lexicon Discovery Using Acoustic Word EmbeddingsPAPERUnsupervised Lexicon Discovery from Acoustic InputPAPERWeakly supervised spoken term discovery using cross-lingual side informationDATACALLHOME Spanish Speech
Speech Synthesis
WIKISpeech synthesisPAPERNatural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram PredictionsPAPERWaveNet: A Generative Model for Raw AudioPAPERTacotron: Towards End-to-End Speech SynthesisPAPERDeep Voice 3: 2000-Speaker Neural Text-to-SpeechPAPEREfficiently Trainable Text-to-Speech System Based on Deep Convolutional Networks with Guided AttentionDATAThe World English BibleDATALJ Speech DatasetDATALessac DataCHALLENGEBlizzard Challenge 2017PRODUCTLyrebirdPROJECTThe Festvox projectTOOLKITMerlin: The Neural Network (NN) based Speech Synthesis System
Speech Enhancement
WIKISpeech enhancementBOOKSpeech enhancement: theory and practicePAPERAn Experimental Study on Speech Enhancement BasedonDeepNeuralNetworkPAPERA Regression Approach to Speech Enhancement BasedonDeepNeuralNetworksPAPERSpeech Enhancement Based on Deep Denoising Autoencoder
Speech-To-Text
- See Automatic Speech Recognition
Spoken Term Detection
- See Speech Segmentation
Stemming
WIKIStemmingPAPERA BACKPROPAGATION NEURAL NETWORK TO IMPROVE ARABIC STEMMINGTOOLKITNLTK Stemmers
Term Extraction
WIKITerminology extractionPAPERNeural Attention Models for Sequence Classification: Analysis and Application to Key Term Extraction and Dialogue Act Detection
Text Similarity
WIKISemantic similarityPAPERA Survey of Text Similarity ApproachesPAPERLearning to Rank Short Text Pairs with Convolutional Deep Neural NetworksPAPERImproved Semantic Representations From Tree-Structured Long Short-Term Memory NetworksCHALLENGESemEval-2014 Task 3: Cross-Level Semantic SimilarityCHALLENGESemEval-2014 Task 10: Multilingual Semantic Textual SimilarityCHALLENGESemEval-2017 Task 1: Semantic Textual SimilarityWIKISemantic Textual Similarity Wiki
Text Simplification
WIKIText simplificationPAPERAligning Sentences from Standard Wikipedia to Simple WikipediaPAPERProblems in Current Text Simplification Research: New Data Can HelpDATANewsela Data
Text-To-Speech
- See Speech Synthesis
Textual Entailment
WIKITextual entailmentPROJECTTextual Entailment with TensorFlowPAPERTextual Entailment with Structured Attentions and CompositionCHALLENGESemEval-2014 Task 1: Evaluation of compositional distributional semantic models on full sentences through semantic relatedness and textual entailmentCHALLENGESemEval-2013 Task 7: The Joint Student Response Analysis and 8th Recognizing Textual Entailment Challenge
Transliteration
WIKITransliterationINFOTransliteration of Non-Latin scriptsPAPERA Deep Learning Approach to Machine TransliterationCHALLENGENEWS 2016 Shared Task on Transliteration of Named EntitiesPROJECTNeural Japanese Transliteration—can you do better than SwiftKey™ Keyboard?
Voice Conversion
PAPERPHONETIC POSTERIORGRAMS FOR MANY-TO-ONE VOICE CONVERSION WITHOUT PARALLEL DATA TRAININGPROJECTDeep neural networks for voice conversion (voice style transfer) in TensorflowPROJECTAn implementation of voice conversion system utilizing phonetic posteriorgramsCHALLENGEVoice Conversion Challenge 2016CHALLENGEVoice Conversion Challenge 2018DATACMU_ARCTIC speech synthesis databasesDATATIMIT Acoustic-Phonetic Continuous Speech Corpus
Voice Recognition
- See Speaker recognition
Word Embeddings
WIKIWord embeddingTOOLKITGensim: word2vecTOOLKITfastTextTOOLKITGloVe: Global Vectors for Word RepresentationINFOWhere to get a pretrained modelPROJECTPre-trained word vectorsPROJECTPre-trained word vectors of 30+ languagesPROJECTPolyglot: Distributed word representations for multilingual NLPPROJECTBPEmb: a collection of pre-trained subword embeddings in 275 languagesCHALLENGESemEval 2018 Task 10 Capturing Discriminative AttributesPAPERBilingual Word Embeddings for Phrase-Based Machine TranslationPAPERA Survey of Cross-Lingual Embedding Models
Word Prediction
INFOWhat is Word Prediction?PAPERThe prediction of character based on recurrent neural network language modelPAPERAn Embedded Deep Learning based Word PredictionPAPEREvaluating Word Prediction: Framing Keystroke SavingsDATAAn Embedded Deep Learning based Word PredictionPROJECTWord Prediction using Convolutional Neural Networks—can you do better than iPhone™ Keyboard?CHALLENGESemEval-2018 Task 2, Multilingual Emoji Prediction
Word Segmentation
WIKIWord segmentationPAPERNeural Word Segmentation Learning for ChinesePROJECTConvolutional neural network for Chinese word segmentationTOOLKITStanford Word SegmenterTOOLKITNLTK Tokenizers
Word Sense Disambiguation
DATAWord-sense disambiguationPAPERTrain-O-Matic: Large-Scale Supervised Word Sense Disambiguation in Multiple Languages without Manual Training DataDATATrain-O-Matic DataDATABabelNet
Related repositories
Hugging Face Transformers (huggingface/transformers) is a Python library that centralizes model definitions for text, computer vision, audio, video, and multimodal machine learning, covering both inference and training. The README describes it as a pivot point compatible with training frameworks such as Axolotl, DeepSpeed, and PyTorch-Lightning, and inference engines such as vLLM, SGLang, and TGI, with more than 1M+ model checkpoints listed on the Hugging Face Hub.
LLMs-from-scratch is Sebastian Raschka's GitHub companion to the Manning book Build a Large Language Model (From Scratch), ISBN 9781633437166. It walks through building, pretraining, and finetuning a GPT-style model in PyTorch across seven chapters and five appendices, plus a growing set of bonus notebooks covering newer open architectures.
Superpowers is an open-source (MIT) skills library and bootstrap instruction set that turns a coding agent's ad-hoc habits into a fixed pipeline: brainstorm a spec, write a plan, build under TDD, review, then close out the branch. Skills trigger automatically once installed, and the same methodology works across eleven different agent harnesses, each requiring its own install step.
TensorFlow is Google's open-source, end-to-end platform for machine learning, hosted at tensorflow/tensorflow under the Apache-2.0 license. It was originally built within Google Brain's Machine Intelligence team for ML and neural network research, and today it ships stable Python and C++ APIs alongside GPU, CPU-only, and Docker install paths. The README positions it as covering both research work and shipping ML-powered applications.
Quick answers
How active is development on Kyubyong/nlp_tasks?
The most recent commit recorded on Kyubyong/nlp_tasks was 7.9 years ago, based on the GitHub push timestamp. The repository has 537 forks — one of the better signals of community interest.
How does Kyubyong/nlp_tasks compare to other AI Tools projects?
Kyubyong/nlp_tasks is tracked by TopGit in the AI Tools category, with 3.0k GitHub stars. Browse the AI Tools topic page on TopGit to compare it against similar projects by stars and activity.
How many stars does Kyubyong/nlp_tasks have?
Kyubyong/nlp_tasks has 3.0k GitHub stars — refresh the page for the live number, or check github.com/Kyubyong/nlp_tasks. TopGit mirrors GitHub's count but does not claim minute-by-minute accuracy.
Is Kyubyong/nlp_tasks open source?
Yes — Kyubyong/nlp_tasks ships under the Apache-2.0 license, which makes its source code freely readable (and, depending on license terms, forkable and reusable). Source: github.com/Kyubyong/nlp_tasks.
What is Kyubyong/nlp_tasks?
Kyubyong/nlp_tasks (Kyubyong/nlp_tasks) is a multi-language project on GitHub. From the project's own README: Natural Language Processing Tasks and References
Where do I read more about Kyubyong/nlp_tasks?
This TopGit page is a snapshot — the READ ME tab shows the project's own README content (links stripped, images preserved). The GitHub repository at github.com/Kyubyong/nlp_tasks is the definitive source.
Read full README in the tab above.
Still deciding about nlp_tasks?
One click hands the question to an AI along with this page — see what it says about nlp_tasks.