This python repo is to learn to scrape data from a website
Be sure to add that lovely star 😀 and fork it for your own copy
- A first try at learning how to scrape a website using python
- It's for anyone who wants to try web scraping in python
Before using this repo fork and clone Link
- follow the instructions in that particular repo for deployment
bs4, requests, pandas, sqlalchemy, psycopg2
downlaod python from python.org
-
add .env file
- pip freeze > requirements.txt pip freeze outputs the package and its version installed in the current environment in the form of a configuration file that can be used with pip install -r.
-Create a
runtime.txtfile to specify the version of python installed for heroku deployment- python(your version of python)
It specifies the python version for all the packages incase there is some descrepency in the packages used
-
Add a
Procfileto specify the commands that are executed by the app on startup. -
You can use a Procfile to declare a variety of process types, including: Your app's web server. Multiple types of worker processes.
-
Deploy the app on heroku following this link
-
Once it is deployed follow these steps to connect to a database
-
To connect python to a node db
- heroku addons:attach <postgres_instance_name> -a <python_app_name>
-
To run python script
- heroku ps:scale web=1
-
To check database has data
-
heroku pg:psql <database_url>
-
select * from historical_data(table name)
-
Checkout the repo PersonalProject} for the next part of the project
Have fun testing and improving it! 😎