High-performance geospatial imagery I/O library supporting NITF, GeoTIFF, JPEG 2000, DTED. Rust core with Python bindings, cloud-native tile access via Zarr, and simple APIs for reading/writing imagery.
Open source. Open possibilities.
Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.
A little curiosity. A world of open source.
THE FIRST COLLECTIONSearXNG is a free, privacy-respecting metasearch engine that aggregates results from multiple search services and databases without tracking or profiling users.
An asynchronous Python client for Stockholm's Open Data Platform API. It currently retrieves disabled-parking locations (about 2045) with address, district, coordinates and validity dates, requires an API key, and is structured so other datasets from the same platform can be added.
A-share full-stack data toolkit with 12-layer architecture, 60 endpoints, 22 data sources, zero authentication, providing packaged A-share data acquisition for AI coding assistants.
SQLModel is a library for interacting with SQL databases in Python, designed to reduce code duplication by combining SQLAlchemy and Pydantic.
An educational repository teaching supervised machine learning from first principles using Python, covering mathematical foundations and practical implementations.
J.A.R.V.I.S is a local-first OSINT platform that uses AI and web scraping to build comprehensive intelligence dossiers on individuals while ensuring data privacy.
rrHog is an open-source, self-hosted web analytics and session replay platform built on rrweb, using FastAPI, NATS JetStream, ClickHouse, Postgres, and Next.js for reliable, async ingestion and fast querying.
Maigret is an OSINT command-line tool that builds a dossier on a person from a username alone, checking 3000+ sites, extracting profile data, and exporting reports in HTML, PDF, JSON, CSV and graph formats.
An automated pipeline that scrapes GitHub Trending, enriches entries via the GitHub API, generates LLM-written summaries and trend analysis, then publishes Markdown reports plus JSON data through a Docusaurus site deployed by GitHub Actions.
Python and MCP example for an Apify actor that returns US House and Senate congressional stock trades and financial disclosures as structured JSON, with filters for member, ticker, and date range.
Taipy is a Python framework designed for data scientists and ML engineers to build and deploy production-ready data and AI-driven web applications without needing to learn front-end languages.
Luigi is a Python package for building complex pipelines of batch jobs, handling dependency resolution, workflow management, visualization, failure handling and command-line integration, with built-in Hadoop support.
Microsoft's 10-week, 20-lesson data science curriculum for beginners, with project-based lessons, quizzes, and 50+ language translations.
LEANN is a lightweight vector database designed for personal RAG applications, reducing storage requirements by up to 97% through graph-based selective recomputation.
deck.gl is a GPU-powered visualization framework designed for high-performance rendering of large-scale data sets using WebGL2 and WebGPU.
CocoIndex is an open-source incremental indexing framework designed to provide AI agents and LLM applications with continuously fresh context by reprocessing only changed data.
FiftyOne is an open-source tool for building high-quality datasets and computer vision models. It enables users to visualize, label, evaluate, and curate visual AI data.
An MCP server for South Korean listed companies' ESG data, created by a securities analyst. It allows AI agents like Claude and ChatGPT to query ratings from 5 agencies, governance indicators, and GHG emissions from KRX, KIND, and GIR via 12 tools without requiring API keys.
Apache Spark is a unified analytics engine designed for large-scale data processing, offering high-level APIs and an optimized engine for general computation graphs.