Automated course data scraper built with Python & Selenium π΅οΈββοΈ
Python Selenium Jupyter NumPy Pandas Kaggle
Course Scraper is a Python-based automation tool that scrapes Coursera courses using Selenium.
As of today (1-Oct-2025), it works smoothly and extracts structured course data into a clean format for analysis.
Why it exists?
I built this to collect educational data for experiments, ML projects, and analytics.
Flexibility:
- If Coursera changes its site paths, you can easily update them inside
_config.py. - Anyone can reuse and modify this project for their own data collection needs.
Warning
Maintenance: This project is no longer being maintained.
Support:
If you run into issues or bugs, just hit me up
You can either download the ZIP or clone via Git:
Download ZIP: Download Here
or
Clone via Git:
git clone https://github.com/cmd-HMN/Course_Scraper.git
cd Course_ScraperYou can run the scraper in two ways: via shell script or directly with Python.
# For Linux, first make the script executable chmod +x ./run_pipeline.sh # Then run the pipeline ./run_pipeline.sh
# Move into source directory
cd scraper/src
# Install dependencies
pip install -r requirements.txt
# Run pipeline
python pipeline.py
Both methods (shell script or Python) accept arguments to control scraping:
_config.pyβ This general setting is in this file.
-co,-overwriteβ Overwrite thecrawler.txtfile (default: False)-cinterval,-crawl_intervalβ Time interval in which the crawler works before taking a rest (default: 5 seconds)-crest,-crawl_restβ Rest time of the crawler after interval (default: 1.5 seconds)
-sinterval,-scraper_intervalβ Time interval in which the scraper works before taking a rest (default: 5 seconds)-srest,-scraper_restβ Rest time of the scraper after interval (default: 1.5 seconds)-scount,-scraper_link_countβ Number of links to scrape (default: all / None)
python pipeline.py -co True -cinterval 6 -crest 2 -sinterval 4 -srest 1 -scount 50
This project is licensed under the MIT License.
See the LICENSE file for more details.
- Built with β€οΈ using Python, Selenium, NumPy, and Pandas
- Inspired by educational data scraping projects
- If you find a bug or issue, just hit me up βοΈ
- π I use Arch btw π
Thank you for checking out Course Scraper!