An automated pipeline that scrapes GitHub Trending, enriches entries via the GitHub API, generates LLM-written summaries and trend analysis, then publishes Markdown reports plus JSON data through a Docusaurus site deployed by GitHub Actions.
Open source. Open possibilities.
Discover quality open-source projects, submit projects anonymously, and claim and edit your own project.
A little curiosity. A world of open source.
THE FIRST COLLECTIONPython and MCP example for an Apify actor that returns US House and Senate congressional stock trades and financial disclosures as structured JSON, with filters for member, ticker, and date range.
Taipy is a Python framework designed for data scientists and ML engineers to build and deploy production-ready data and AI-driven web applications without needing to learn front-end languages.
Luigi is a Python package for building complex pipelines of batch jobs, handling dependency resolution, workflow management, visualization, failure handling and command-line integration, with built-in Hadoop support.
Microsoft's 10-week, 20-lesson data science curriculum for beginners, with project-based lessons, quizzes, and 50+ language translations.
LEANN is a lightweight vector database designed for personal RAG applications, reducing storage requirements by up to 97% through graph-based selective recomputation.
deck.gl is a GPU-powered visualization framework designed for high-performance rendering of large-scale data sets using WebGL2 and WebGPU.
CocoIndex is an open-source incremental indexing framework designed to provide AI agents and LLM applications with continuously fresh context by reprocessing only changed data.
FiftyOne is an open-source tool for building high-quality datasets and computer vision models. It enables users to visualize, label, evaluate, and curate visual AI data.
An MCP server for South Korean listed companies' ESG data, created by a securities analyst. It allows AI agents like Claude and ChatGPT to query ratings from 5 agencies, governance indicators, and GHG emissions from KRX, KIND, and GIR via 12 tools without requiring API keys.
Apache Spark is a unified analytics engine designed for large-scale data processing, offering high-level APIs and an optimized engine for general computation graphs.
newspaper3k is a Python 3 library designed for extracting and curating articles from the web, providing tools for full-text and metadata extraction.
Deck Lab is a Python/Flask workspace for building, researching and goldfish-simulating competitive Commander (MTG) decks, with versioned data contracts, isolated QA, and verified container releases.
A local, read‑only dashboard that visualizes Claude Code usage—token counts, cost estimates, sessions, tools, skills, projects, and activity patterns—by scanning ~/.claude/projects/ and serving a FastAPI‑based UI.
PyMC is a Python package for Bayesian statistical modeling featuring advanced MCMC and variational inference algorithms, with intuitive model specification syntax and support for complex probabilistic models.
MCP server exposing Swiss federal geodata (maps, elevation, geocoding, STAC, WMTS, OEREB) via 20 no‑auth tools for AI assistants.
An interactive analytics and data visualization component designed for large, real-time, and streaming datasets.
A CLI tool for processing railway speed restriction data from spreadsheets and visualizing them on HTML maps.
plotly.py is an interactive, open-source, browser-based graphing library for Python, built on plotly.js. It provides 30+ chart types, works in Jupyter notebooks, standalone HTML, and Dash apps, and is MIT licensed.
Licitarium is a Windows program that mirrors PNCP data offline: tenders, contracts, price registrations, PCA plans, and unit prices in a searchable SQLite database, with automated PCA drafting and eight printed reports.