Finite state and Constraint Grammar based morphological analysers, proofing tools, and language resources for the Komi-Zyrian language, part of the GiellaLT project.
Open source. Open possibilities.
Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.
A little curiosity. A world of open source.
THE FIRST COLLECTIONLangExtract is a Python library from Google that uses LLMs to extract structured information from unstructured text, grounding every extraction to its exact source location and generating interactive HTML visualizations.
TextBlob is a Python library for natural language processing, offering a simple API for sentiment analysis, part-of-speech tagging, noun phrase extraction, classification, tokenization, and more, built on NLTK and pattern.
Batchalign3 is a TalkBank audio and ML pipeline for producing and enriching CHAT transcripts, covering ASR, forced alignment, neural morphotagging, translation and utterance segmentation. It ships a CLI, Python package, PyO3 bridge, React dashboard and an experimental Tauri desktop shell.
Free, open-source AI engineering curriculum from first principles: 523 lessons in 20 phases (~342 hours), with math-to-production coverage in Python, TypeScript, Rust, and Julia. Every lesson ships a reusable prompt, skill, agent, or MCP server; includes an installable AI tutor and Claude certification prep.
A pure-Python scientific computing platform with 22 modules covering quantum, ML, statistics, ODE solvers, and more—zero runtime dependencies, with optional NumPy/SciPy adapters for validation and reproducibility.
Haqumei is a Japanese Grapheme-to-Phoneme (G2P) library written in Rust with Python bindings, providing accurate text-to-phoneme conversion with prosody information, word-phoneme mapping, and a CLI tool for speech synthesis frontends.
SpellKit is a fast, safe Ruby gem with a Rust SymSpell implementation. It offers sub-millisecond latency, term protection via regex, hot-reloadable dictionaries, and zero dependencies. It supports multiple instances, skip patterns, and is production-ready for search-term extraction.
Swift Tokenizers is a high-performance Swift wrapper around Hugging Face's Rust tokenizers crate, focused solely on tokenization without Hub dependencies. It supports macOS, iOS and Linux, offering encoding, decoding, streaming detokenization, chat templates and tool calling.
HanLP is a production-ready multilingual NLP toolkit based on PyTorch and TensorFlow 2.x, supporting tokenization, POS tagging, named entity recognition, syntactic parsing, semantic analysis, and more, with both RESTful and native APIs.
Premove ITN is an open-source, context-aware inverse text normalization tool for English voice-agent transcripts. It converts spoken ASR output into canonical written forms using a generate-score-decode pipeline with DeBERTa contextual scoring.
Finite state and Constraint Grammar based morphological analysers, proofing tools, and language resources for the Tsuut'ina (Sarsi) language, part of the GiellaLT open-source linguistic infrastructure.
ModelScope is an open-source Python library implementing Model-as-a-Service, offering unified pipeline, Trainer and MsDataset interfaces for inference, fine-tuning and evaluation of CV, NLP, audio, multi-modal and scientific models, plus model/dataset hub integration.
Hugging Face Transformers is a model-definition framework for machine learning across text, vision, audio, and multimodal tasks. It provides a unified API for inference and training with over 1 million pretrained checkpoints.