141 Commits
Author SHA1 Message Date
timberjoegithub c8296da736 Merge branch '3.0' of https://github.com/timberjoegithub/GoogleScrape into 3.0
# Conflicts:
#	social.py
2026-05-27 13:22:03 +00:00
timberjoegithub 13ae301667 Refactor: repo hygiene, shared utils, focused tests, and CI improvements
- Remove dead imports/duplication from Google2xls.py and social.py
- Centralize shared logic in review_workflow_utils.py
- Add focused posting-flow tests in tests/test_social.py (all pass)
- Update .gitignore for secrets and local artifacts
- Trim requirements.txt to direct deps
- Improve CI: compile all scripts, run focused tests
- Prep repo for clean commit (no secrets, no local artifacts)
2026-05-27 13:17:56 +00:00
Joe Steele 3e400170b0 linting 2024-11-03 12:32:21 -05:00
Joe Steele e6aa6b6339 fixed posting to wordpress and addiition to database 2024-09-13 18:06:43 -04:00
Joe Steele eae7276147 fixed city state tags 2024-08-04 11:40:14 -04:00
Joe Steele 552927bdb6 small edits 2024-07-08 08:21:09 -04:00
Joe Steele 6773143e4f Huge work on database and wordpress updates 2024-07-08 08:01:21 -04:00
Joe Steele 73bb8564ca got montage renamed, uploaded and attached 2024-07-06 08:01:11 -04:00
Joe Steele 0e2d39ef64 working on tiktok and threads 2024-07-03 08:41:21 -04:00
Joe Steele b99e43f0d6 tiktok 2024-07-02 16:25:12 -04:00
Joe Steele 972edffd22 started tiktok 2024-07-02 15:58:20 -04:00
Joe Steele 7110b92de3 merge 2024-07-02 10:34:36 -04:00
Joe Steele f8b6200168 fixed instagram over 60 seconds 2024-07-02 09:03:11 -04:00
Joe Steele 99d7ac8d63 trying to sort of xls still 2024-07-02 06:55:46 -04:00
Joe Steele 80b506ff04 fixing xls 2024-07-01 09:05:22 -04:00
Joe Steele ac9f671cb2 commiting 2024-06-30 21:52:17 -04:00
Joe Steele c6c14afc6b Updated but still broken writing to excel 2024-06-30 21:48:06 -04:00
Joe Steele d774a207a0 Merge branch '3.0' of https://github.com/timberjoegithub/GoogleScrape into 3.0 2024-06-28 13:56:50 -04:00
Joe Steele 6e5d89cdb1 update 2024-06-28 13:56:42 -04:00
Joe Steele d8019d5565 resub 2024-06-28 08:58:25 -04:00
Joe Steele c1a1341c49 update excel 2024-06-27 13:19:32 -04:00
Joe Steele 663b88ab8b dedupr 2024-06-27 08:52:30 -04:00
Joe Steele 2ae3339a63 fixed getting all visit date into same format 2024-06-26 18:25:07 -04:00
Joe Steele 5dd64d5d5e Merge branch '3.0' of https://github.com/timberjoegithub/GoogleScrape into 3.0 2024-06-26 16:11:34 -04:00
Joe Steele 8ffb9f46a8 resync 2024-06-26 16:11:28 -04:00
Joe Steele 2b33ff35c8 fix visit date (seems to work) 2024-06-26 16:02:27 -04:00
Joe Steele 8a9c3e0b76 added test for visited date starting with string 2024-06-26 15:35:36 -04:00
Joe Steele a26b334571 cleanup 2024-06-26 12:24:21 -04:00
Joe Steele 70b7b4ae80 Merge branch '3.0' of https://github.com/timberjoegithub/GoogleScrape into 3.0 2024-06-25 18:12:09 -04:00
Joe Steele 50612bbc02 fix dimensions of xls sheet 2024-06-25 18:11:52 -04:00
Joe Steele 50d391b9e4 linting 2024-06-25 16:29:31 -04:00
Joe Steele 08e08f4b02 more linting 2024-06-25 15:39:30 -04:00
Joe Steele 6b9ab27651 linting 2024-06-25 10:21:44 -04:00
Joe Steele 357bad0e1d Merge branch '3.0' of https://github.com/timberjoegithub/GoogleScrape into 3.0 2024-06-25 09:15:25 -04:00
Joe Steele c6af16b186 reupping 2024-06-25 09:14:51 -04:00
Joe Steele f7067ade72 working on date information for posts with no pics 2024-06-25 08:18:55 -04:00
Joe Steele 7992b28036 added visitdate 2024-06-25 01:09:42 -04:00
Joe Steele 3c44e940e2 started rewriting post to wordpress 2024-06-24 09:41:25 -04:00
Joe Steele 6361f21f8a changed one more instance of timeout 2024-06-21 08:24:11 -04:00
Joe Steele 59b9d2ec89 Merge branch '3.0' of https://github.com/timberjoegithub/GoogleScrape into 3.0 2024-06-20 21:12:39 -04:00
Joe Steele a774e538b5 bringing database up to date 2024-06-20 21:11:12 -04:00
Joe Steele 2c6051a35a more linting 2024-06-20 14:48:00 -04:00
Joe Steele 42b1fe50d4 changed request timeout to be globally configurable in env file 2024-06-20 11:55:16 -04:00
Joe Steele ebd30a9979 added ability to force wordpress update for all posts 2024-06-20 07:59:26 -04:00
Joe Steele 4b271afd7b change timeout for wordpress 2024-06-20 07:01:10 -04:00
Joe Steele 9e48f67e57 merged codbase and wrote some new database writing 2024-06-19 22:47:02 -04:00
Joe Steele 7ec229de9f Merge branch '3.0' of https://github.com/timberjoegithub/GoogleScrape into 3.0 2024-06-19 22:18:35 -04:00
Joe Steele 30b7ce183b more updates 2024-06-19 22:18:30 -04:00
Joe Steele fddb2656db fixed xls rows 2024-06-19 22:12:39 -04:00
Joe Steele ce980774ab working tiktok 2024-06-19 22:10:11 -04:00
Joe Steele 239c3a5434 working on xls transformation 2024-06-19 08:02:43 -04:00
Joe Steele a8c86b45ec Merge branch '3.0' of https://github.com/timberjoegithub/GoogleScrape into 3.0 2024-06-18 17:39:44 -04:00
Joe Steele 37f4e61840 test 2024-06-18 17:39:16 -04:00
Joe SteeleandGitHub e299389d65 Update pylint.yml 2024-06-18 15:36:05 -04:00
Joe Steele 8b14969869 databse writes 2024-06-18 12:59:54 -04:00
Joe Steele 41cd6ead52 Merge branch '3.0' of https://github.com/timberjoegithub/GoogleScrape into 3.0 2024-06-18 08:01:08 -04:00
Joe Steele 6ba0adfa61 working on writing to DB and excel 2024-06-18 08:01:02 -04:00
Joe SteeleandGitHub 4e7475733e Update pylint.yml 2024-06-17 23:30:20 -04:00
Joe Steele 9cf9fd3697 Merge branch '3.0' of https://github.com/timberjoegithub/GoogleScrape into 3.0 2024-06-17 23:27:33 -04:00
Joe Steele a08d22efc6 more linting 2024-06-17 23:27:22 -04:00
Joe SteeleandGitHub 0659dfbd22 Update pylint.yml 2024-06-17 23:19:14 -04:00
Joe Steele e0d3857163 more linting 2024-06-17 23:14:15 -04:00
Joe Steele 5df6f6bd2c fixing stuff that broke in reorganization 2024-06-17 20:04:27 -04:00
Joe Steele 55dfec90df more linting 2024-06-17 17:27:39 -04:00
Joe Steele 058477d87f formatting for linting 2024-06-17 16:57:48 -04:00
Joe Steele 9f480185cd added docstrings 2024-06-17 15:06:34 -04:00
Joe Steele 5385488e01 futher pushed to DB for config 2024-06-17 14:33:58 -04:00
Joe Steele b35e9e093d more updates 2024-06-17 09:27:06 -04:00
Joe Steele 2c16c348c7 cleaning up processing reviews 2024-06-16 23:10:44 -04:00
Joe Steele de67bb2e91 cutting down functions 2024-06-16 22:51:42 -04:00
Joe Steele 3e05188aea fixed socialpost function to work for everything 2024-06-16 22:07:20 -04:00
Joe Steele df0dce5b50 Updated trues to 1's 2024-06-16 17:19:33 -04:00
Joe Steele 9cd8c248dd Fork 3.0 2024-06-15 17:16:54 -04:00
Joe Steele c6924ce3e5 cleanup - linting 2024-06-15 08:15:14 -04:00
Joe Steele cb5757ae13 upfates 2024-06-14 23:09:51 -04:00
Joe Steele 7f25a9de4d twitter working now (although it skips invalid 2024-06-14 16:23:08 -04:00
Joe Steele 0d6aac15f3 fixed twitter 2024-06-14 15:53:10 -04:00
Joe Steele c06d35e0e6 fixing some wordpress posting problems 2024-06-14 13:59:15 -04:00
Joe Steele 274b7698c6 Adding links to Facebook and cleaning up twitter 2024-06-14 13:42:15 -04:00
Joe Steele 49b2f22efc update 2024-06-14 10:14:16 -04:00
Joe Steele efe726aed9 fixed twitter 2024-06-14 08:49:47 -04:00
Joe Steele 697e59f260 fixed more twitter post problems 2024-06-13 23:16:00 -04:00
Joe Steele 4e44c93f64 update Exceptions 2024-06-13 22:05:24 -04:00
Joe Steele ad5ab9ce48 Merge branch 'main' of https://github.com/timberjoegithub/GoogleScrape 2024-06-13 21:26:04 -04:00
Joe Steele 24975dbc38 update 2024-06-13 21:25:52 -04:00
timberjoegithub 2bc34775c0 dint know 2024-06-13 15:38:25 -04:00
Joe Steele 59cb54c6c3 fixed twitter 2024-06-13 11:58:31 -04:00
Joe Steele 9dd2920ff2 removed link to wordpress site from being posted to wordpress site 2024-06-13 08:59:16 -04:00
Joe Steele 9dac335681 Added link to website to wordpress 2024-06-13 08:57:57 -04:00
Joe Steele cbc21642b1 Added business url to Twitter 2024-06-13 08:54:23 -04:00
Joe Steele e2e6175452 Merge branch 'main' of https://github.com/timberjoegithub/GoogleScrape 2024-06-13 08:51:08 -04:00
Joe Steele dfb0d4b6f6 added link to wordpress review to Twitter 2024-06-13 08:49:42 -04:00
Joe Steele d4c10e064d more comments 2024-06-12 22:18:47 -04:00
Joe Steele eadbdbf648 Merge branch 'main' of https://github.com/timberjoegithub/GoogleScrape 2024-06-12 22:07:46 -04:00
Joe Steele 30a501edb8 added comments 2024-06-12 22:06:53 -04:00
Joe Steele 8b42fda122 more docstrings 2024-06-12 14:30:23 -04:00
Joe Steele 62fc5307f3 Updated docstrings 2024-06-12 14:24:44 -04:00
Joe Steele c753dbcd86 Linting 2024-06-12 14:01:36 -04:00
Joe Steele 16c3c75ddc added requirements 2024-06-12 13:15:39 -04:00
Joe Steele 8f2897bcb3 add new lines 2024-06-12 12:25:51 -04:00
Joe Steele 53350d0c1d more formatting 2024-06-12 09:48:53 -04:00
Joe Steele ee0110eae9 working on wpurl, need to clean DB, got dup records 2024-06-11 17:33:52 -04:00
Joe Steele 09d694d937 Merge branch 'main' of https://github.com/timberjoegithub/GoogleScrape 2024-06-11 15:49:33 -04:00
Joe Steele b69bb36c24 working on wpurl 2024-06-11 15:47:19 -04:00
Joe Steele 3e9a2b9fc3 worked on making wpurl write correctly to DB 2024-06-11 15:41:55 -04:00
Joe Steele c91c1b1465 move inactive code to side 2024-06-11 13:25:30 -04:00
Joe Steele fe23ebba62 Moved all inactive code to the side 2024-06-11 13:24:57 -04:00
Joe Steele 0b33c745e7 Cleaning up more 2024-06-11 10:29:04 -04:00
Joe Steele f16290a0d0 Fixed database updates 2024-06-10 21:20:47 -04:00
Joe Steele ce825d3ddb Fixed problems with wordpress adds to database 2024-06-10 20:29:53 -04:00
Joe Steele 1c8f2af062 oo 2024-06-10 15:32:51 -04:00
Joe Steele 83ce83f3f8 more cleaning 2024-06-10 15:14:14 -04:00
Joe Steele caace8be66 Major cleanup and moving code around. Likely to break something 2024-06-10 14:48:51 -04:00
Joe Steele e6273febe5 Had some string compares against ints 2024-06-10 12:12:41 -04:00
Joe Steele f3794f3472 new fixing 2024-06-10 11:02:01 -04:00
Joe Steele 9590a3ef28 added comments 2024-06-09 19:26:13 -04:00
Joe Steele fb22ae2707 Merge branch 'main' of https://github.com/timberjoegithub/GoogleScrape 2024-06-09 19:23:46 -04:00
Joe Steele db5878aca4 see previous 2024-06-09 19:21:49 -04:00
Joe Steele d65dd91841 Move updates to database and add new columns for cool stuff 2024-06-09 19:21:42 -04:00
timberjoegithub 765e651116 Update 2024-06-07 14:19:35 -04:00
Joe Steele 8d4fa44a1b Moved accounting to database 2024-06-07 12:34:09 -04:00
Joe Steele 88ae7df0eb Merge branch 'main' of https://github.com/timberjoegithub/GoogleScrape 2024-06-07 10:49:04 -04:00
Joe Steele 6a01354521 successfully writing database status for weg 2024-06-07 08:15:42 -04:00
timberjoegithub 40bbc1b37d fixed database 2024-06-07 07:02:00 -04:00
Joe Steele 368da2609e added columns for different socials 2024-06-07 06:36:27 -04:00
Joe Steele fe703c9c65 update database 2024-06-06 17:33:51 -04:00
Joe Steele 484ce73ac2 Merge branch 'main' of https://github.com/timberjoegithub/GoogleScrape 2024-06-05 09:10:06 -04:00
Joe Steele 5f33cfc912 resync 2024-06-05 09:09:39 -04:00
Joe Steele 6aa64288c5 more linting 2024-06-05 08:32:55 -04:00
Joe Steele 4def1bc002 more linting 2024-06-05 08:30:45 -04:00
Joe Steele ad3645fbe0 more linting 2024-06-05 08:26:22 -04:00
Joe Steele edf0678a49 fixing linting problems 2024-06-05 08:00:09 -04:00
Joe Steele 62d70040c8 Merge branch 'main' of https://github.com/timberjoegithub/GoogleScrape 2024-06-04 08:29:41 -04:00
Joe Steele f434214eff UPDATE 2024-06-04 08:26:34 -04:00
Joe Steele 467f22c1b6 jj 2024-06-03 14:41:07 -04:00
Joe Steele 53f7ae3c7e Twitter v2 updates 2024-06-03 12:51:48 -04:00
Joe Steele bd2491c406 fixing xtwitter 2024-06-02 21:38:49 -04:00
timberjoegithub 68dba21de7 initial add of X/twitter 2024-06-02 18:59:22 -04:00
timberjoegithub 9b26c7ccb7 cv 2024-05-30 11:41:56 -04:00
timberjoegithub 3f6133db0d out of fync 2024-05-30 11:40:43 -04:00
Joe Steele 6996d9b9e5 resync 2024-05-29 10:05:07 -04:00
10 changed files with 5278 additions and 929 deletions
+10 -7
View File
@@ -1,23 +1,26 @@
name: Pylint name: Validate
on: [push] on: [push]
jobs: jobs:
build: validate:
runs-on: ubuntu-latest runs-on: ubuntu-latest
strategy: strategy:
matrix: matrix:
python-version: ["3.8", "3.9", "3.10"] python-version: ["3.11"]
steps: steps:
- uses: actions/checkout@v3 - uses: actions/checkout@v3
- name: Set up Python ${{ matrix.python-version }} - name: Set up Python ${{ matrix.python-version }}
uses: actions/setup-python@v3 uses: actions/setup-python@v3
with: with:
python-version: ${{ matrix.python-version }} python-version: ${{ matrix.python-version }}
- name: Install dependencies - name: Install validation dependencies
run: | run: |
python -m pip install --upgrade pip python -m pip install --upgrade pip
pip install pylint pip install python-dateutil selenium
- name: Analysing the code with pylint - name: Compile core scripts
run: | run: |
pylint $(git ls-files '*.py') python -m py_compile Google2xls.py WP2xls.py social.py text.py review_workflow_utils.py
- name: Run focused unit tests
run: |
python -m unittest discover -s tests
+1 -13
View File
@@ -3,19 +3,7 @@ __pycache__/
out.xlsx out.xlsx
Output/ Output/
env.py env.py
requirements.txt
test.py
ThreadsPost.py
TikTokPost.py
WP2xls.py
Xpost.py
Instapost.py
InstaBot.py
Google2xls.py
google.py
geckodriver.log geckodriver.log
Facepost.py
config/ config/
env.py
social.py
.vscode/ .vscode/
debug.log
+152
View File
@@ -0,0 +1,152 @@
import time
import os
from selenium import webdriver
from selenium.webdriver.chrome.webdriver import WebDriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By
import re
import urllib3
from urllib.request import urlretrieve
from openpyxl import Workbook
import pandas as pd
from env import URL, DriverLocation
from datetime import datetime
from review_workflow_utils import (
build_google_chrome_options,
build_review_media_path,
count_google_review_pages,
normalize_google_media_url,
sanitize_review_name,
scroll_google_reviews,
set_stealth_driver,
)
today = datetime.today().strftime('%Y-%m-%d')
# Rotate User Agent would be helpful
def get_data(driver):
"""
this function get main text, score, name
"""
print('get data...')
# Click on more botton on each text reviews
more_elemets = driver.find_elements(By.CSS_SELECTOR, '.w8nwRe.kyuRq')
for list_more_element in more_elemets:
list_more_element.click()
# Find Pictures that have the expansion indicator to see the rest of the pictures under them and click it to expose them all
more_pics = driver.find_elements(By.CLASS_NAME, 'Tya61d')
for list_more_pics in more_pics:
if 'showMorePhotos' in list_more_pics.get_attribute("jsaction") :
print('Found extra pics')
list_more_pics.click()
elements = driver.find_elements(By.CLASS_NAME, 'jftiEf')
lst_data = []
for data in elements:
name = data.find_element(By.CSS_SELECTOR, 'div.d4r55.YJxk2d').text
try: address = data.find_element(By.CSS_SELECTOR, 'div.RfnDt.xJVozb').text
except: address = 'Unknonwn'
print ('Name of location: ',name, ' Address:',address)
try: visitdate = data.find_element(By.CSS_SELECTOR, 'span.rsqaWe').text
except: visitdate = "Unknown"
print('Visited: ',visitdate)
try: text = data.find_element(By.CSS_SELECTOR, 'div.MyEned').text
except: text = ''
try: score = data.find_element(By.CSS_SELECTOR, 'span.kvMYJc').get_attribute("aria-label") #find_element(By.CSS_SELECTOR,'aria-label').text #) ##QA0Szd > div > div > div.w6VYqd > div:nth-child(2) > div > div.e07Vkf.kA9KIf > div > div > div.m6QErb.DxyBCb.kA9KIf.dS8AEf > div.m6QErb > div:nth-child(3) > div:nth-child(2) > div > div:nth-child(4) > div.DU9Pgb > span.kvMYJc
except: score = "Unknown"
more_specific_pics = data.find_elements(By.CLASS_NAME, 'Tya61d')
pics= []
pics2 = []
# check to see if folder for pictures and videos already exists, if not, create it
cleanname = sanitize_review_name(name)
if not os.path.exists('./Output/Pics/'+cleanname):
os.makedirs('./Output/Pics/'+cleanname)
# Walk through all the pictures and videos for a given review
for lmpics in more_specific_pics:
# Grab URL from style definiton (long multivalue string), and remove the -p-k so that it is full size
urlmedia = normalize_google_media_url(lmpics.get_attribute("style"))
print ('URL : ',urlmedia)
pics.append(urlmedia)
# time.sleep(2)
# photoindex = str(lmpics.get_attribute("data-photo-index"))
# Grab the name of the file and remove all spaces and special charecters to name the folder
filename = re.sub( r'[^a-zA-Z0-9]','', str(lmpics.get_attribute("aria-label")))
# filename = re.sub( r'[^a-zA-Z0-9]','', filename)
if lmpics == more_specific_pics[0]:
lmpics.click()
time.sleep(2)
#iframe = driver.find_element(By.TAG_NAME, "iframe")
tempdate = str((driver.find_element(By.CLASS_NAME,'mqX5ad')).text).rsplit("-",1)
visitdate = re.sub( r'[^a-zA-Z0-9]','',tempdate[1])
print ('Visited: ',visitdate)
# driver.switch_to.default_content()
# Check to see if it has a sub div, which represents the label with the video length displayed, this will be done
# because videos are represented by pictures in the main dialogue, so we need to click through and grab the video URL
if (lmpics.find_elements(By.CSS_SELECTOR,'div.fontLabelMedium.e5A3N')) :
ext='.mp4'
lmpics.click()
time.sleep(2)
# After we click the right side is rendered in an inframe, Store iframe web element
iframe = driver.find_element(By.TAG_NAME, "iframe")
# switch to selected iframe
driver.switch_to.frame(iframe)
# Now find button and click on button
video_elements = driver.find_elements(By.XPATH ,'//video') #.get_attribute('src')
urlmedia = str((video_elements[0]).get_attribute("src"))
# return back away from iframe
driver.switch_to.default_content()
else:
# The default path if it is not a video link
ext='.jpg'
# Add the correct extension to the file name
filename = filename+ext
# Test to see if file already exists, and if it does not grab the media and store it in location folder
picsLocalpath = build_review_media_path(name, visitdate, filename)
if not os.path.isfile(picsLocalpath):
urlretrieve(urlmedia, picsLocalpath)
# Store the local path to be used in the excel document
pics2.append(picsLocalpath)
dictPostComplete= {'google':1,'web':0,'yelp':0,'facebook':0,'xtwitter':0,'Instagram':0,'tiktok':0}
lst_data.append([name , text, score,pics,pics2,"GoogleMaps",visitdate,address,dictPostComplete])
return lst_data
# Grab a count of how far we need to scroll
def counter(driver):
return count_google_review_pages(driver)
# Do the scrolling
def scrolling(driver, review_count):
return scroll_google_reviews(driver, review_count, label='scrolling...')
def write_to_xlsx(data):
print('write to excel...')
cols = ["name", "comment", 'rating','picsURL','picsLocalpath','source','date','address','dictPostComplete']
df = pd.DataFrame(data, columns=cols)
df.to_excel('./Output/reviews.xlsx')
if __name__ == "__main__":
print('starting...')
options = build_google_chrome_options(
ignore_certificate_errors=True,
ignore_ssl_errors=True,
)
# Setting the driver path and requesting a page
driver = webdriver.Chrome(options=options) # Firefox(options=options)
# Changing the property of the navigator value for webdriver to undefined
set_stealth_driver(driver)
driver.get(URL)
time.sleep(5)
review_count = counter(driver)
scrolling(driver, review_count)
data = get_data(driver)
driver.close()
write_to_xlsx(data)
print('Done!')
+283
View File
@@ -0,0 +1,283 @@
import ast
import base64
import pandas as pd
import requests
from openpyxl import Workbook, load_workbook
from env import wpAPOurl, xls,user, password,wpAPI
import datetime as dt
import json
import re
#import socket
import urllib3
from datetime import datetime
from review_workflow_utils import normalize_wordpress_date
today = datetime.today().strftime('%Y-%m-%d')
#print(urllib.request.urlopen("https://www.stackoverflow.com").getcode())
def is_port_open(host, port):
# sock = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
try:
# removehttps = re.sub(r'https://','', host)
# removetrailing = removehttps.split('/')[0]
isWebUp = urllib3.request("GET", host)
#sock.close()
# result = sock.connect_ex((host, port))
if isWebUp.status == 200:
return True
except Exception as error:
print ('Could not open port to website: ', host, type(error))
return False
# def to_json(obj):
# return json.dumps(obj, default=lambda obj: obj,__dict__ )
def connect_and_auth_wp( headers):
# connect and auth
data_string = f"{user}:{password}"
token = base64.b64encode(data_string.encode()).decode("utf-8")
headers = {"Authorization": f"Basic {token}"}
return (headers)
def check_media(filename, headers):
# Regex gilename to format like in WordPress media name
#file_name_minus_extension = re.sub(r'\'|(....$)','', filename, flags=re.IGNORECASE)
file_name_minus_extension = filename
response = requests.get(wpAPI + "/media?search="+file_name_minus_extension, headers=headers)
try:
result = response.json()
file_id = int(result[0]['id'])
link = result[0]['guid']['rendered']
return file_id, link
except Exception:
print(' No existing media with same name in Wordpress media folder: ' + filename)
return False, False
def check_post(postname,postdate, headers):
response = requests.get(wpAPI+"/posts?search="+postname, headers=headers)
try:
result = response.json()
post_id = int(result[0]['id'])
post_date = result[0]['date']
if postdate == post_date:
return post_id
except Exception:
print('No existing post with same name: ' + postname)
return False
def post_to_wp(title, content, headers,date, rating,address, picslist):
# post
NewPost = False
countreview = False
addresshtml = re.sub(" ", ".",address)
googleadress = r"<a href=https://www.google.com/maps/dir/?api=1&destination="+addresshtml + r">"+address+r"</a>" # https://www.google.com/maps/dir/?api=1&destination=760+West+Genesee+Street+Syracuse+NY+13204
contentpics = ''
picl = picslist[1:-1]
pic2 = picl.replace(",","")#re.sub(r',','',picl) #re.sub( r'[^a-zA-Z0-9]','',tempdate[1])
pic3= pic2.replace("'","")
# print (pic3)
pidchop = pic3.split(" ")
linkslist=[]
# linkslist.clear
print (' Figuring out date of Post : ',title)
date = normalize_wordpress_date(date)
try:
post_id = check_post(title, str(date), headers)
except Exception as error:
print ('Could not check to see post already exists', type(error).__name__)
post_id = False
if ( post_id == False):
# if ((title) != False):
googleadress = r"<a href=https://www.google.com/maps/dir/?api=1&destination="+addresshtml + r">"+address+r"</a>"
post_data = {
"title": title,
# "content": address+'\n\n'+content+'\n'+rating+'\n\n' ,
"content": googleadress+'\n\n'+content+'\n'+rating ,
"status": "publish", # Set to 'draft' if you want to save as a draft
"date": date,
# "date": str(newdate)+'T22:00:00',
# "author":"joesteele"
}
try:
# response = requests.post(wpAPOurl, json=post_data, headers=headers)
# myobj = {'json': post_data, 'headers': headers }
# response = requests.post(wpAPOurl, json = myobj)
headers2 = headers
response = requests.post(wpAPOurl, json = post_data, headers=headers2)
if ( response.status_code != 201 ):
print ('Error: ',response)
# print ('json:'+ post_data+' headers :'+headers)
# else:
# for loop in linkslist:
# print (' Adding ', loop['link'], ' to posting')
# try:
# contentpics += r'<img src="'+ loop['link'] + r' alt="' + title +r'">' +'\n\n'
# except Exception as error:
# print("An error occurred:", type(error).__name__) # An error occurred:
# post_data = ()
# response2 = requests.post(wpAPI+"/media/" + response['id'], json = post_data, headers=headers2)
else:
NewPost = True
post_id_json = response.json()
post_id = post_id_json.get('id')
print (' New post is has post_id = ',post_id)
except Exception as error:
print("An error occurred:", type(error).__name__) # An error occurred:
postneedsupdate = True
else:
print (' Post already existed: Post ID : ',post_id)
for pic in pidchop:
picslice2 = pic.split("/")[-1]
picslice = picslice2.split(".")
picname = picslice[0]
caption =title
description = title+"\n"+address
print (' Found Picture: ',picname)
file_id, link = check_media(picname, headers)
# link = linknew['rendered']
if (file_id) is False:
print (' '+picname+' was not already found in library, adding it')
countreview = True
image = {
"file": open(pic, "rb"),
"post": post_id,
"caption": caption,
"description": description
}
try:
image_response = requests.post(wpAPI + "/media", headers=headers, files=image)
except Exception as error:
print(" An error uploading picture ' + picname+ ' occurred:", type(error).__name__) # An error occurred:
if ( image_response.status_code != 201 ):
print (' Error- Image ',picname,' was not successfully uploaded. response: ',image_response)
else:
PicDict=image_response.json()
file_id= PicDict.get('id')
link = PicDict.get('guid').get("rendered")
print (' ',picname,' was successfully uploaded to website with ID: ',file_id, link)
try:
linksDict = {'file_id' : file_id , 'link' : link}
linkslist.append(linksDict)
except Exception as error:
print(" An error adding to dictionary " , file_id , link , " occurred:", type(error).__name__) # An error occurred:
else:
print (' Photo ',picname,' was already in library and added to post with ID: ',file_id,' : ',link)
try:
image_response = requests.post(wpAPI + "/media/" + str(file_id), headers=headers, data={"post" : post_id})
except Exception as error:
print (' Error- Image ',picname,' was not attached to post. response: ',image_response)
try:
post_response = requests.get(wpAPI + "/posts/" + str(post_id), headers=headers)
if (link in str(post_response.text)):
print (' Image link for ', picname, 'already in content of post: ',post_id, post_response.text, link)
else:
linkslist.append({'file_id' : file_id , 'link' : link})
countreview = True
except Exception as error:
print(" An error loading the metadata from the post " + post_response.title + ' occurred:", type(error).__name__) # An error occurred')
#ratinghtml = post_response.text
firstMP4 = True
for piclink in linkslist:
#for loop in linkslist:
print (' Adding ', piclink['link'], ' to posting')
try:
ext = (piclink['link'].split( '.')[-1] )
if (ext == 'mp4'):
if (firstMP4):
contentpics += '\n' +r'[evp_embed_video url="' + piclink['link'] + r'" autoplay="true"]'
firstMP4 = False
else:
contentpics += '\n' +r'[evp_embed_video url="' + piclink['link'] + r'"]'
#[evp_embed_video url="http://example.com/wp-content/uploads/videos/vid1.mp4" autoplay="true"]
else:
contentpics += '\n '+r'<div class="col-xs-4"><img id="'+str(file_id)+r'"' + r'src="' + piclink['link'] + r'"></div>'
# contentpics += '\n '+r'<img src="'+ piclink['link'] + '> \n'
#contentpics += r'<img src="'+ piclink['link'] + r' alt="' + title +r'">' +'\n\n'
except Exception as error:
print("An error occurred:", type(error).__name__) # An error occurred:
try:
response_piclinks = requests.post(wpAPI+"/posts/"+ str(post_id), data={"content" : googleadress+'\n\n'+content+'\n'+rating + contentpics, "featured_media" : file_id}, headers=headers)
except Exception as error:
print(" An error writing images to the post " + post_response.title + ' occurred:", type(error).__name__) # An error occurred')
return NewPost
# def add_image(wpAPI,pic, picname,caption, description,headers):
# # loop thru
# image = {
# "file": open(pic, "rb"),
# "caption": caption,
# "description": description
# }
# try:
# image_response = requests.post(wpAPI + "/media", headers=headers, files=image)
# except Exception as error:
# print("An error occurred:", type(error).__name__) # An error occurred:
# if ( image_response != 201 ):
# print (image_response)
# # print ('json:'+ post_data+' headers :'+headers)
# print ('image response: ',image_response)
# return image_response
# def read_from_xlsx(ws):
# print('read from excel...')
# return (ws)
def process_reviews(ws, headers):
# Process
needreversed = False
totalcount = 2
count = 0
rows = list(ws.iter_rows(min_row=1, max_row=ws.max_row))
if needreversed:
rows = reversed(rows)
for processrow in rows:
if (count >= totalcount):
print ('Exceeded the number of posts per run, exiting')
break
# for processrow in (ws.rows):
# If reviewxls (comment has content) and (picsURL has connect)
if processrow[1].value != "name": # Skip header line of xls sheet
print ("Processing : ",processrow[1].value)
# ast.literal_eval(deployments) = processrow[9].value
writtento = (ast.literal_eval(processrow[9].value))
# Check to see if the website has already been written to according to the xls sheet, if it has not... then process
if ((writtento["web"]) == 0) and (is_port_open(wpAPI, 443)) and (processrow[2].value != None) :
try:
NewPost = post_to_wp(processrow[1].value, processrow[2].value, headers ,processrow[7].value, processrow[3].value, processrow[8].value, processrow[5].value)
try:
writtento["web"] = 1
#processrow[9] = writtento[9]
# processrow[9].value = '"'+writtento+'"'
processrow[9].value = str(writtento)
except Exception as error:
print("An error occurred writing value Excel file:", type(error).__name__) # An error occurred:
print ('Success Posting: '+processrow[1].value)# ',processrow[1].value, processrow[2].value, headers,processrow[7].value, processrow[3].value,processrow[8].value, processrow[5].value, temp3["web"] )
if NewPost == True:
count +=1
try:
wb.save(xls)
except Exception as error:
print("An error occurred writing Excel file:", type(error).__name__) # An error occurred:
except:
print ('Error writing post : ',processrow[1].value, processrow[2].value, headers,processrow[7].value, processrow[3].value,processrow[8].value, processrow[5].value, writtento["web"] )
return (headers)
if __name__ == "__main__":
countreview = False
headers = {}
wb = load_workbook(filename = xls)
ws = wb['Sheet1']
print('starting...')
# read_from_xlsx(ws)
print('Connect and auth to Wordpress')
data_string = f"{user}:{password}"
token = base64.b64encode(data_string.encode()).decode("utf-8")
headers = {"Authorization": f"Basic {token}"}
# print ('Headers:',headers)
connect_and_auth_wp( headers)
print('Processing Reviews')
# print ('Headers:',headers)
process_reviews(ws, headers)
print('Done!')
+14 -13
View File
@@ -1,13 +1,14 @@
autopep8==1.5.7 aiohttp
et-xmlfile==1.1.0 googlemaps
numpy==1.21.2 instagrapi
openpyxl==3.0.9 jsonpickle
pandas==1.3.3 moviepy
pkg_resources==0.0.0 mysqlclient
pycodestyle==2.7.0 openpyxl
python-dateutil==2.8.2 pandas
pytz==2021.1 python-dateutil
selenium==3.141.0 requests
six==1.16.0 selenium
toml==0.10.2 sqlalchemy
urllib3==1.26.6 tweepy
urllib3
+121
View File
@@ -0,0 +1,121 @@
import re
import time
from datetime import datetime, timedelta
from pathlib import Path
from dateutil.relativedelta import relativedelta
GOOGLE_REVIEW_SCROLL_XPATH = (
'//*[@id="QA0Szd"]/div/div/div[1]/div[2]/div/div[1]/div/div/div[5]/div[2]'
)
GOOGLE_REVIEW_SCROLL_SCRIPT = (
'document.getElementsByClassName("dS8AEf")[0].scrollTop = '
'document.getElementsByClassName("dS8AEf")[0].scrollHeight'
)
def build_google_chrome_options(
headless=False,
log_level=None,
ignore_certificate_errors=False,
ignore_certificate_errors_spki_list=False,
ignore_ssl_errors=False,
remote_debugging_pipe=False,
):
from selenium import webdriver
options = webdriver.ChromeOptions()
if log_level is not None:
options.add_argument(f"--log-level={log_level}")
if ignore_certificate_errors:
options.add_argument("--ignore-certificate-error")
if ignore_certificate_errors_spki_list:
options.add_argument("--ignore-certificate-errors-spki-list")
if ignore_ssl_errors:
options.add_argument("--ignore-ssl-errors")
if headless:
options.add_argument("--headless")
options.add_argument("--lang=en-US")
options.add_argument("--disable-blink-features=AutomationControlled")
if remote_debugging_pipe:
options.add_argument("--remote-debugging-pipe")
options.add_experimental_option("excludeSwitches", ["enable-automation"])
options.add_experimental_option("useAutomationExtension", False)
return options
def set_stealth_driver(driver):
driver.execute_script(
"Object.defineProperty(navigator, 'webdriver', {get: () => undefined})"
)
def count_google_review_pages(driver):
result = driver.find_element("class name", 'Qha3nb').text
result = result.replace(',', '')
result = result.split(' ')
result = result[0].split('\n')
return int(result[0]) // 10 + 1
def scroll_google_reviews(driver, review_count, label='google_scroll...', delay_seconds=3):
print(label, end="")
time.sleep(delay_seconds)
scrollable_div = driver.find_element("xpath", GOOGLE_REVIEW_SCROLL_XPATH)
google_scroller = None
for _ in range(review_count):
try:
google_scroller = driver.execute_script(
GOOGLE_REVIEW_SCROLL_SCRIPT,
scrollable_div,
)
time.sleep(delay_seconds)
print('.', end="")
except Exception as error:
print(f"Error while {label.strip('.')}: {error}")
break
print('')
return google_scroller
def sanitize_review_name(name):
return re.sub(r'[^a-zA-Z0-9]', '', name)
def normalize_google_media_url(style_text):
return re.sub(r'=\S*-p-k-no', '=-no', (re.findall(r"['\"](.*?)['\"]", style_text))[0])
def build_review_media_path(review_name, visit_date, filename, base_dir="./Output/Pics"):
review_dir = Path(base_dir) / sanitize_review_name(review_name) / visit_date
review_dir.mkdir(parents=True, exist_ok=True)
return str(review_dir / filename)
def normalize_wordpress_date(date_text, now=None):
reference_time = now or datetime.now()
normalized = str(date_text).strip()
lowered = normalized.lower()
if lowered == "a day ago":
publish_at = reference_time - timedelta(days=1)
elif "day" in lowered:
publish_at = reference_time - timedelta(days=int(re.sub(r'[^0-9]', '', lowered) or 0))
elif lowered == "a week ago":
publish_at = reference_time - relativedelta(weeks=1)
elif "week" in lowered:
publish_at = reference_time - relativedelta(weeks=int(re.sub(r'[^0-9]', '', lowered) or 0))
elif lowered == "a month ago":
publish_at = reference_time - relativedelta(months=1)
elif "month" in lowered:
publish_at = reference_time - relativedelta(months=int(re.sub(r'[^0-9]', '', lowered) or 0))
elif lowered == "a year ago":
publish_at = reference_time - relativedelta(years=1)
elif "year" in lowered:
publish_at = reference_time - relativedelta(years=int(re.sub(r'[^0-9]', '', lowered) or 0))
else:
compact_date = normalized.replace(" ", "")
publish_at = datetime.strptime(compact_date[:3] + "/" + compact_date[3:] + "/01", '%b/%Y/%d')
return publish_at.strftime('%Y-%m-%dT22:00:00')
+4505 -838
View File
File diff suppressed because it is too large Load Diff
+62
View File
@@ -0,0 +1,62 @@
import unittest
from datetime import datetime
from review_workflow_utils import (
build_review_media_path,
count_google_review_pages,
normalize_google_media_url,
normalize_wordpress_date,
sanitize_review_name,
)
class FakeElement:
def __init__(self, text):
self.text = text
class FakeDriver:
def __init__(self, text):
self.text = text
def find_element(self, by, value):
self.last_lookup = (by, value)
return FakeElement(self.text)
class ReviewWorkflowUtilsTests(unittest.TestCase):
def test_count_google_review_pages_rounds_up_by_review_page(self):
driver = FakeDriver("25 reviews")
page_count = count_google_review_pages(driver)
self.assertEqual(page_count, 3)
self.assertEqual(driver.last_lookup, ("class name", "Qha3nb"))
def test_normalize_wordpress_date_handles_relative_weeks(self):
publish_date = normalize_wordpress_date("2 weeks ago", now=datetime(2026, 5, 27, 9, 30, 0))
self.assertEqual(publish_date, "2026-05-13T22:00:00")
def test_normalize_wordpress_date_handles_month_year_strings(self):
publish_date = normalize_wordpress_date("Jan 2025")
self.assertEqual(publish_date, "2025-01-01T22:00:00")
def test_normalize_google_media_url_strips_thumbnail_suffix(self):
style_text = 'background-image: url("https://example.com/photo=s1600-p-k-no");'
self.assertEqual(
normalize_google_media_url(style_text),
"https://example.com/photo=-no",
)
def test_build_review_media_path_sanitizes_review_name(self):
media_path = build_review_media_path("Joe's Diner!", "20260527", "front.jpg", base_dir="/tmp/review-tests")
self.assertEqual(media_path, "/tmp/review-tests/JoesDiner/20260527/front.jpg")
self.assertEqual(sanitize_review_name("Joe's Diner!"), "JoesDiner")
if __name__ == "__main__":
unittest.main()
+50
View File
@@ -0,0 +1,50 @@
import unittest
from unittest.mock import patch, MagicMock
import social
class SocialPostingFlowTests(unittest.TestCase):
def setUp(self):
# Patch env and Posts for isolation
self.env_patch = patch('social.env')
self.mock_env = self.env_patch.start()
self.mock_env.request_timeout = 1
self.mock_env.facebooksleep = 0
self.mock_env.forcegoogleupdate = False
self.mock_env.xls = 'dummy.xlsx'
self.mock_env.mariadb = False
self.mock_env.mariadbuser = 'user'
self.mock_env.mariadbpass = 'pass'
self.mock_env.mariadbserver = 'localhost'
self.mock_env.mariadbdb = 'db'
self.mock_env.needreversed = False
self.mock_env.postssession = MagicMock()
self.mock_env.posts = []
self.mock_env.xlsdf = MagicMock()
self.mock_env.posts = []
self.mock_env.postssession.query.return_value.count.return_value = 0
self.addCleanup(self.env_patch.stop)
@patch('social.requests.post')
@patch('builtins.open', new_callable=MagicMock)
def test_post_facebook_video_handles_success(self, mock_open, mock_post):
mock_post.return_value.json.return_value = {'id': '123'}
mock_open.return_value.__enter__.return_value = MagicMock()
result = social.post_facebook_video('groupid', ['video.mp4'], 'token', 'title', 'content', 'date', 'rating', 'address')
self.assertEqual(result, {'id': '123'})
@patch('social.requests.post', side_effect=AttributeError('fail'))
@patch('builtins.open', new_callable=MagicMock)
def test_post_facebook_video_handles_error(self, mock_open, mock_post):
mock_open.return_value.__enter__.return_value = MagicMock()
result = social.post_facebook_video('groupid', ['video.mp4'], 'token', 'title', 'content', 'date', 'rating', 'address')
self.assertFalse(result)
def test_write_to_database_handles_empty(self):
# Should not raise
data = []
local_outputs = {'xlsdf': MagicMock(), 'posts': [], 'postssession': MagicMock()}
result = social.write_to_database(data, local_outputs)
self.assertEqual(result, data)
if __name__ == '__main__':
unittest.main()
+25 -3
View File
@@ -1,3 +1,14 @@
"""
Sends a text message to a phone number through the specified carrier.
Args:
phone_number_inside (str): The phone number to send the message to.
carrier_inside (str): The carrier of the phone number.
message_inside (str): The message content to be sent.
Returns:
None
"""
import smtplib import smtplib
import sys import sys
@@ -11,13 +22,24 @@ CARRIERS = {
EMAIL = "EMAIL" EMAIL = "EMAIL"
PASSWORD = "PASSWORD" PASSWORD = "PASSWORD"
def send_message(phone_number, carrier, message): def send_message(phone_number_inside, carrier_inside, message_inside):
recipient = phone_number + CARRIERS[carrier] """
Sends a text message to a phone number through the specified carrier.
Args:
phone_number_inside (str): The phone number to send the message to.
carrier_inside (str): The carrier of the phone number.
message_inside (str): The message content to be sent.
Returns:
None
"""
recipient = phone_number_inside + CARRIERS[carrier_inside]
auth = (EMAIL, PASSWORD) auth = (EMAIL, PASSWORD)
server = smtplib.SMTP("smtp.gmail.com", 587) server = smtplib.SMTP("smtp.gmail.com", 587)
server.starttls() server.starttls()
server.login(auth[0], auth[1]) server.login(auth[0], auth[1])
server.sendmail(auth[0], recipient, message) server.sendmail(auth[0], recipient, message_inside)
if __name__ == "__main__": if __name__ == "__main__":
if len(sys.argv) < 4: if len(sys.argv) < 4: