OPEN SOURCE, OPEN TO EVERYONE

Open source. Open possibilities.

Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.

Human-curated · Discover open source3941discovered

A little curiosity. A world of open source.

THE FIRST COLLECTION
swisstopo-mcpmalkreide
ADDED

MCP server exposing Swiss federal geodata (maps, elevation, geocoding, STAC, WMTS, OEREB) via 20 no‑auth tools for AI assistants.

Developer toolsData & databasesBackend & APIs
dots.ocrstudio-dots-ai
ADDED

dots.ocr is a multilingual document layout parsing vision-language model that extracts text, tables, formulas and layout from images and PDFs, and can convert charts and diagrams into SVG code.

AI & MLData & databasesMultimodal AI
unionunionlabs
ADDED

A zero-knowledge interoperability layer for cross-chain messaging, asset transfers, NFTs, and DeFi, using consensus verification and IBC, with nodes, provers, relayers, and SDKs.

Self-hostedData & databasesServer management
ADDED

A library of 166 ready-to-use Agent Skills plus 100+ scientific database integrations that turn AI coding agents into research assistants for biology, chemistry, medicine, drug discovery and data analysis.

AI & MLData & databasesAI agents
kalauzgy-mate
ADDED

A CLI tool for processing railway speed restriction data from spreadsheets and visualizing them on HTML maps.

Data & databasesDeveloper toolsData integration
StormByte-BufferStormBytePP
ADDED

StormByte-Buffer is a C++26 library providing FIFO, SharedFIFO, Ring buffers, and pipeline tools for the StormByte suite. Includes thread-safe/non-thread-safe buffers, producer/consumer patterns, and configurable pipeline stages (sync/async/parallel). Requires StormByte Base.

Developer toolsData & databasesComponent libraries
seatunnelapache
ADDED

Apache SeaTunnel is a high-performance, distributed data integration tool supporting 160+ connectors, batch-stream integration, and multi-engine execution.

Data & databasesData integration
beamapache
ADDED

Apache Beam is a unified programming model for defining both batch and streaming data-parallel processing pipelines.

Data & databasesData integration
mediacrawlerNanmiCoder
ADDED

MediaCrawler is an open-source multi-platform social media data collection tool supporting Xiaohongshu, Douyin, Kuaishou, Bilibili, Weibo, Tieba, Zhihu, and more. It enables keyword search, post ID crawling, secondary comment extraction, creator profile scraping, login state caching, IP proxy pool, and comment word cloud generation, with data export in CSV, JSON, Excel, SQLite, and MySQL formats.

Data & databasesAutomationData integration
cliPortabase
ADDED

Portabase CLI is an official command-line tool for managing Portabase instances, supporting database backup/restore for PostgreSQL, MySQL, MongoDB, Redis, and more, with Docker volume handling.

Developer toolsData & databasesCLI & terminal
python-liegeklaasnicolaas
ADDED

An asynchronous Python client for the Open Data Platform of Liège (Belgium). It currently exposes the disabled parking spaces and garages datasets through a simple async API, and its code base is designed to be extended to other datasets from the same platform.

Developer toolsData & databasesBackend & APIs
toolboxGlatzel
ADDED

A collection of low-level, reusable Python and Rust utilities designed as foundational building blocks for systems development and data processing.

Developer toolsData & databasesComponent libraries
data-inclusiongip-inclusion
ADDED

data·inclusion is an open-source project that aggregates, cleans, geocodes, and publishes open data on social and professional integration in France, sourced from public and partner databases.

Data & databasesSelf-hostedData integration
pydanticpydantic
ADDED

Pydantic is a data validation library for Python that leverages type hints to enforce data structures.

Developer toolsData & databasesBackend & APIs
kefcoremasesgroup
ADDED

KEFCore is an Entity Framework Core provider for Apache Kafka, enabling .NET applications to use Kafka topics as a distributed database. It offers full LINQ support and allows interaction with Kafka through standard EF Core models, abstracting away Kafka-specific client code.

Developer toolsData & databasesBackend & APIs

A self-hosted data API for Douyin and TikTok. It fetches posts, profiles, comments, and playlists, and downloads videos and image albums without watermarks. Features a self-healing identity pool, PostgreSQL archiving, REST API, MCP server, CLI, and a web console.

Self-hostedData & databasesServer management
rustycsvjeffhuen
ADDED

RustyCSV is a high-performance CSV parser for Elixir, built as a Rust NIF with SIMD acceleration, parallel parsing, and bounded-memory streaming. It serves as a drop-in replacement for NimbleCSV.

Developer toolsData & databasesBackend & APIs
market-data-collectorGoldenmaplefezez5520
ADDED

FlashLoanArbitrage is a Node.js bot and smart contract for DeFi flash-loan arbitrage. The README says goflash.js monitors ETH/USDC prices on Chainlink, Uniswap V2, SushiSwap, Curve and Balancer, waits for a 0.9% spread, and triggers a contract that borrows USDC, trades ETH, repays the loan and converts profit back to ETH.

AutomationData & databasesBots & integrations
doclingdocling-project
ADDED

Docling parses diverse document formats including PDF, DOCX, PPTX, XLSX, HTML, EPUB, images, audio, and more, with advanced PDF understanding, export to Markdown/JSON, OCR, VLM support, and integrations with LangChain, LlamaIndex, and other AI frameworks for RAG and agentic workflows.

AI & MLData & databasesRAG & knowledge
ADDED

DataPulse is an open, read-only trust and verification layer for Malaysian public data. It repeatedly probes 418 official datasets, publishes machine-readable evidence on freshness, licence, schema and provenance, and exposes the catalogue to AI agents through a 19-tool read-only MCP server.

AI & MLData & databasesAI agents
Load more projects