Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
345 commits
Select commit Hold shift + click to select a range
5cde25a
updates
AseelOmer Jul 2, 2025
72edfec
updates
AseelOmer Jul 3, 2025
fedaf46
updates
AseelOmer Jul 3, 2025
4c14e7f
trying to check if mermaid diagrams work
Alaa-Elgozouli Jul 4, 2025
91061d2
added a graph to the data_preparation readme
Alaa-Elgozouli Jul 4, 2025
44edf30
created a file for checking the most frequent words in the fake posts
Alaa-Elgozouli Jul 4, 2025
6437596
fixing linting errors
Alaa-Elgozouli Jul 4, 2025
d74e0e8
fixing linting issues
Alaa-Elgozouli Jul 4, 2025
05159ae
fixing linting errors
Alaa-Elgozouli Jul 4, 2025
377a3a1
fixing linting errors
Alaa-Elgozouli Jul 4, 2025
2a9de52
fixing linting errors
Alaa-Elgozouli Jul 4, 2025
c4d2737
fixing linting errors
Alaa-Elgozouli Jul 4, 2025
2557f9e
fixing linting errors
Alaa-Elgozouli Jul 4, 2025
ab54ea9
fixing linting errors
Alaa-Elgozouli Jul 4, 2025
2c08ba0
Format notebook to pass Ruff check
Alaa-Elgozouli Jul 4, 2025
e179ed2
Separating lines to pass linting errors
Alaa-Elgozouli Jul 4, 2025
6c2350c
Separating lines to pass linting errors
Alaa-Elgozouli Jul 4, 2025
375bb9b
ughhh linting erros again
Alaa-Elgozouli Jul 4, 2025
93ef577
fixing flake8 and ruff format
Alaa-Elgozouli Jul 5, 2025
289328a
added code for visualization
Alaa-Elgozouli Jul 5, 2025
5dfd635
added csv file and visual for most frequently used words in fake posts
Alaa-Elgozouli Jul 5, 2025
eec49d3
Merge pull request #44 from MIT-Emerging-Talent/fake_most_frequent_words
AseelOmer Jul 5, 2025
1f2ee74
Merge pull request #45 from MIT-Emerging-Talent/1_datasets
AseelOmer Jul 5, 2025
e2af846
edited the path file for both cvs and png files for most frequent wor…
Alaa-Elgozouli Jul 5, 2025
73f812d
update CSV and vscode settings
Alaa-Elgozouli Jul 5, 2025
d39ef02
--
Alaa-Elgozouli Jul 5, 2025
ab2e334
Exploration Data
majdadel20 Jul 5, 2025
8478aa1
Merge branch 'main' of https://github.com/MIT-Emerging-Talent/ET6-CDS…
majdadel20 Jul 5, 2025
61ad7dc
cleaned the real jobs in the Aegean dataset for comparsion and analys…
Alaa-Elgozouli Jul 6, 2025
d2e44d9
fixed linting issues
Alaa-Elgozouli Jul 7, 2025
f7430f8
added a brief description of the notebook
Alaa-Elgozouli Jul 7, 2025
74972ec
added a missing word
Alaa-Elgozouli Jul 7, 2025
5d01f5a
changed the fake most used words file path to exploration folder
Alaa-Elgozouli Jul 7, 2025
ed795ab
changing fake words file to exploartion folder
Alaa-Elgozouli Jul 7, 2025
6711521
fixing formatting errors
Alaa-Elgozouli Jul 7, 2025
562afc5
Merge pull request #48 from MIT-Emerging-Talent/real_jobs_extraction_…
Alaa-Elgozouli Jul 8, 2025
03dd8e3
added a section for checking the most frequent words from the real ra…
Alaa-Elgozouli Jul 8, 2025
fbabf0d
added a code for checking the most frequent words in LLM-refined vers…
Alaa-Elgozouli Jul 11, 2025
4852465
Fix Ruff formatting in most_frequent_words.ipynb
Alaa-Elgozouli Jul 11, 2025
8e03763
Update Exploration File
majdadel20 Jul 13, 2025
87fd363
Merge branch 'main' of https://github.com/MIT-Emerging-Talent/ET6-CDS…
majdadel20 Jul 13, 2025
f1315ad
Update after review for exploration readme
majdadel20 Jul 13, 2025
cebbc03
changes
AseelOmer Jul 14, 2025
c351e9b
added few comments for better understanding in the most frequent word…
Alaa-Elgozouli Jul 14, 2025
adddfbf
fixed formatting errors
Alaa-Elgozouli Jul 14, 2025
5c127a8
cleaned the arranging of the folders and the 1_datasets folder
Alaa-Elgozouli Jul 14, 2025
0685859
fixing linting errors
Alaa-Elgozouli Jul 14, 2025
28fb2c5
organized 1_datasets folder
Alaa-Elgozouli Jul 14, 2025
bc4a77b
Update Exploration File
majdadel20 Jul 15, 2025
81e7512
Merge branch 'main' of https://github.com/MIT-Emerging-Talent/ET6-CDS…
majdadel20 Jul 15, 2025
0899037
added new fakejobs refinements steps
Elocodes Jul 16, 2025
d30d7da
Merge branch 'main' into data_preprocessing
Elocodes Jul 16, 2025
2ce65e8
fixed formatting errors
Elocodes Jul 16, 2025
de8d8ca
Merge pull request #56 from MIT-Emerging-Talent/data_preprocessing
AseelOmer Jul 16, 2025
9a78ba8
update
AseelOmer Jul 16, 2025
2b466f7
Revert "update"
AseelOmer Jul 16, 2025
c93b08f
Refined fake jobs batch - Aseel
AseelOmer Jul 16, 2025
70e01bd
Merge pull request #59 from MIT-Emerging-Talent/refine-84-fake-jobs-a…
Elocodes Jul 16, 2025
11fc3e0
updated the cleaned data folder in 1_datasets. you'll find everything…
Alaa-Elgozouli Jul 17, 2025
027f835
updated the 1_datasets folder
Alaa-Elgozouli Jul 17, 2025
3e55655
Merge pull request #63 from MIT-Emerging-Talent/updated_cleaned_data
AseelOmer Jul 17, 2025
1489456
worked on data cleaning
Alaa-Elgozouli Jul 17, 2025
1aaa729
data refining
AseelOmer Jul 18, 2025
7c6dbad
chore: adjust flake8 max line length to 120
AseelOmer Jul 18, 2025
46c9feb
data refining
AseelOmer Jul 18, 2025
728cbc7
Fix formatting to pass Ruff lint
AseelOmer Jul 18, 2025
af4fece
Format notebook to comply with Ruff
AseelOmer Jul 18, 2025
f42cac0
Fix line-length linting errors (E501) in notebook
AseelOmer Jul 18, 2025
e285a2f
Fix linting issues with nbqa and ruff
AseelOmer Jul 18, 2025
c8ea507
Final E501 fix: shorten long print line
AseelOmer Jul 18, 2025
e566382
Fix ruff formatting for CI
AseelOmer Jul 18, 2025
bb30884
Fix ruff formatting for CI
AseelOmer Jul 18, 2025
f485d9b
Analyze urgency tone using sentiment polarity across job post types (…
Rouaa93 Jul 18, 2025
9ddb768
Merge pull request #64 from MIT-Emerging-Talent/data3
Alaa-Elgozouli Jul 18, 2025
c364bd9
fix formatting error
Rouaa93 Jul 18, 2025
efd3bc0
fix linting error
Rouaa93 Jul 18, 2025
9a7af93
generated more llm-refined posts and fixed lint errors
Alaa-Elgozouli Jul 18, 2025
5068434
refined fake jobs
AseelOmer Jul 18, 2025
ca1e3af
Merge pull request #66 from MIT-Emerging-Talent/refine_84_fake_jobs_majd
Alaa-Elgozouli Jul 18, 2025
0123c7e
refined jobs
AseelOmer Jul 18, 2025
172dd04
Merge pull request #67 from MIT-Emerging-Talent/refine-data-84
Rouaa93 Jul 18, 2025
4535e57
refined
AseelOmer Jul 18, 2025
a82c4cf
Merge pull request #68 from MIT-Emerging-Talent/refining-data
Elocodes Jul 18, 2025
6d242e9
Merge pull request #65 from MIT-Emerging-Talent/urgency-tone
Elocodes Jul 18, 2025
cab41b1
added a section for cleaning the data in the llm refinement code
Alaa-Elgozouli Jul 19, 2025
ef272de
worked on refinement code
Alaa-Elgozouli Jul 19, 2025
c08ba95
Merge branch 'main' of https://github.com/MIT-Emerging-Talent/ET6-CDS…
Alaa-Elgozouli Jul 19, 2025
9a6865f
worked on the code
Alaa-Elgozouli Jul 19, 2025
6c2dc0b
fixed ci checks errors
Alaa-Elgozouli Jul 19, 2025
aaa1c96
added more refined posts in cleaned data
Alaa-Elgozouli Jul 19, 2025
ad5446f
new defined fake jobs_Geehan
GeehanAli Jul 19, 2025
7287b18
Geehan New
GeehanAli Jul 19, 2025
5d1e823
llm-refine
AseelOmer Jul 20, 2025
bda0a38
Merge pull request #71 from MIT-Emerging-Talent/llm-refine
Alaa-Elgozouli Jul 20, 2025
d31cc8e
added more llm refined posts
Alaa-Elgozouli Jul 20, 2025
a26f30b
Accepted remote version during merge
Alaa-Elgozouli Jul 20, 2025
2254bd6
added more llm refined posts
Alaa-Elgozouli Jul 20, 2025
e4c6a87
Merge pull request #72 from MIT-Emerging-Talent/llm-refine
AseelOmer Jul 20, 2025
1559fd4
llm-refine
AseelOmer Jul 20, 2025
4096c1c
Merge pull request #73 from MIT-Emerging-Talent/llm-refine
Alaa-Elgozouli Jul 20, 2025
0af0532
Merge branch 'main' of https://github.com/MIT-Emerging-Talent/ET6-CDS…
Alaa-Elgozouli Jul 20, 2025
0310c18
TF-IDF
majdadel20 Jul 21, 2025
4043e19
add top fake job chart for TFIDF
majdadel20 Jul 21, 2025
3bd1e5b
fixing linting error
majdadel20 Jul 21, 2025
3b8e03e
fixing formatting issues
GeehanAli Jul 21, 2025
5f5e4e2
chore: ignore .vscode folder in ls-lint
GeehanAli Jul 21, 2025
ac66668
fixing checks errors
GeehanAli Jul 21, 2025
663853b
fixing checks errors
GeehanAli Jul 21, 2025
cc22fa6
Hope this will work fixing cheks errors
GeehanAli Jul 21, 2025
bbafd9d
github you're driving me crazy
GeehanAli Jul 21, 2025
a04449a
Clean N-gram notebook with no unrelated changes
AseelOmer Jul 21, 2025
76ca99e
well I'm about to give up
GeehanAli Jul 21, 2025
e83462b
Fix lint: shorten long lines, import display function
AseelOmer Jul 21, 2025
8d30e23
n-gram
AseelOmer Jul 21, 2025
7be9e65
n-gram
AseelOmer Jul 21, 2025
92a148b
ngram
AseelOmer Jul 21, 2025
5ee04c1
ngram
AseelOmer Jul 21, 2025
717c0c2
working full analysis pipeline
Elocodes Jul 21, 2025
8251dfa
organized datasets folder and working on full analysis pipeline
Elocodes Jul 21, 2025
4c16327
Merge branch 'main' into data_preprocessing
Elocodes Jul 21, 2025
64c8fbb
Merge pull request #70 from MIT-Emerging-Talent/Refined_84_jobs_Geehan
AseelOmer Jul 21, 2025
a5bd074
add retrospective m3
Rouaa93 Jul 21, 2025
2d645b7
Merge pull request #79 from MIT-Emerging-Talent/really-clean-n-gram
Rouaa93 Jul 21, 2025
217eb4f
added Geehan's note
GeehanAli Jul 22, 2025
ca1fd32
Update 5_data_analysis.md
GeehanAli Jul 22, 2025
06882a8
sending in exploration and analysis of real and fake jobs
Elocodes Jul 22, 2025
1e091ce
Merge branch 'main' into data_preprocessing
Elocodes Jul 22, 2025
6647546
resolved conflict that arose by a file name change - fake_jobs_AIrefi…
Elocodes Jul 22, 2025
ca42a19
fixed formatting issues
Elocodes Jul 22, 2025
10f4654
fixed formatting issue in the recombine file and renamed aegean/Hypat…
Elocodes Jul 22, 2025
3dc2df8
fixed formatting issue
Elocodes Jul 22, 2025
e46c53b
Fix: Corrected folder and file casing for aegean500_vs_hypatia500_dat…
Elocodes Jul 22, 2025
296afae
Merge pull request #74 from MIT-Emerging-Talent/TFIDF
Elocodes Jul 22, 2025
b78747d
Merge pull request #85 from MIT-Emerging-Talent/data_preprocessing
AseelOmer Jul 22, 2025
3d631e3
readability
AseelOmer Jul 22, 2025
189b28e
add Geehan's retrospective m3
Rouaa93 Jul 22, 2025
c8d8a72
add Geehan's retrospective= m3
Rouaa93 Jul 22, 2025
287eb20
readability
AseelOmer Jul 22, 2025
01792fc
readability
AseelOmer Jul 22, 2025
8b731a7
Merge pull request #86 from MIT-Emerging-Talent/readability
Elocodes Jul 22, 2025
2b98035
finalized words frequency code
Alaa-Elgozouli Jul 22, 2025
e184b81
fixing lint errors
Alaa-Elgozouli Jul 22, 2025
dd09036
add Justina's retrospective= m3
Rouaa93 Jul 22, 2025
1c8d4ab
Merge branch 'main' of https://github.com/MIT-Emerging-Talent/ET6-CDS…
Alaa-Elgozouli Jul 22, 2025
192d2a1
edit Justina's retrospective= m3
Rouaa93 Jul 22, 2025
95126d3
edit Justina's- retrospective= m3
Rouaa93 Jul 22, 2025
027a2e9
edit Justina's.. retrospective= m3
Rouaa93 Jul 22, 2025
e42dad2
edit Justina's\ retrospective= m3
Rouaa93 Jul 22, 2025
8a74b0b
Merge pull request #87 from MIT-Emerging-Talent/all_most_frequent_words
AseelOmer Jul 22, 2025
afdc327
edit Alaa's\ retrospective= m3
Rouaa93 Jul 22, 2025
869433b
Merge branch 'main' of https://github.com/MIT-Emerging-Talent/ET6-CDS…
Alaa-Elgozouli Jul 22, 2025
fbe1dad
completed the departments and analysis comparison code
Alaa-Elgozouli Jul 22, 2025
71f5c8e
fixed format erros
Alaa-Elgozouli Jul 22, 2025
389a03e
fixed lint errors
Alaa-Elgozouli Jul 22, 2025
1745d26
added a title for the all sections graph
Alaa-Elgozouli Jul 22, 2025
83cafa6
added a title in the all sections graph
Alaa-Elgozouli Jul 22, 2025
99c09b2
restore notebook to last known good commit 1559fd4
Alaa-Elgozouli Jul 22, 2025
9f4fd4a
Merge pull request #91 from MIT-Emerging-Talent/fix-llm-refinement
AseelOmer Jul 23, 2025
11201c5
readme
AseelOmer Jul 22, 2025
03b18a2
update
AseelOmer Jul 23, 2025
eee50fa
readme
AseelOmer Jul 23, 2025
a1543b0
readme
AseelOmer Jul 23, 2025
23a3d97
Merge pull request #94 from MIT-Emerging-Talent/readme-milestone3
Alaa-Elgozouli Jul 23, 2025
38edcdb
Merge pull request #81 from MIT-Emerging-Talent/add-retrospective-m3
Alaa-Elgozouli Jul 23, 2025
ae6eb7e
update
AseelOmer Jul 24, 2025
1f33512
update
AseelOmer Jul 24, 2025
1bf86a8
Merge pull request #90 from MIT-Emerging-Talent/analysis_departments_…
AseelOmer Jul 24, 2025
4d74ffb
enhancement
AseelOmer Jul 25, 2025
9db0207
updates
AseelOmer Jul 25, 2025
b97d7a0
update
GeehanAli Jul 25, 2025
cd5b919
updates
AseelOmer Jul 25, 2025
03289cf
updates
AseelOmer Jul 25, 2025
626a1e3
Remove old notebooks to fix CI
AseelOmer Jul 25, 2025
b31b4ba
Merge remote-tracking branch 'origin/main' into readability
AseelOmer Jul 25, 2025
d657547
Fix CI: Clean up flake8 formatting issues in notebooks
AseelOmer Jul 25, 2025
ae9aed2
deleted last update
GeehanAli Jul 26, 2025
bfa643d
update
AseelOmer Jul 26, 2025
080bfac
updates
AseelOmer Jul 26, 2025
072e80e
excluded company profile from the words count analysis
Alaa-Elgozouli Jul 26, 2025
7b667ac
commiting the visuals as well
Alaa-Elgozouli Jul 26, 2025
6b7dd9c
updates
AseelOmer Jul 26, 2025
339a428
resolving merge conflicts
Alaa-Elgozouli Jul 26, 2025
86e2ae0
updates
AseelOmer Jul 26, 2025
59f08ae
updates
AseelOmer Jul 26, 2025
489a4ff
updates
AseelOmer Jul 26, 2025
db1e100
updates
AseelOmer Jul 26, 2025
f276216
Data Analysis Readme
majdadel20 Jul 26, 2025
a152883
Merge pull request #101 from MIT-Emerging-Talent/all_most_frequent_words
AseelOmer Jul 26, 2025
dfd28b8
Merge pull request #100 from MIT-Emerging-Talent/readability
Alaa-Elgozouli Jul 26, 2025
0fdba32
updates
AseelOmer Jul 27, 2025
b60bdf0
updates
AseelOmer Jul 27, 2025
2f96e4c
Resolved all conflicts using our version in readme-m3
AseelOmer Jul 27, 2025
fbda582
updates
AseelOmer Jul 27, 2025
1f89776
Merge pull request #108 from MIT-Emerging-Talent/readme-m3
majdadel20 Jul 27, 2025
bd30661
worked on and updated the data preparation readme
Alaa-Elgozouli Jul 27, 2025
5f503e4
worked on and updated the data preparation readme
Alaa-Elgozouli Jul 27, 2025
34627bd
worked on and updated the datasets readme
Alaa-Elgozouli Jul 27, 2025
6cb0765
Merge pull request #109 from MIT-Emerging-Talent/updated_data_prepara…
AseelOmer Jul 27, 2025
1c3be9b
Merge pull request #110 from MIT-Emerging-Talent/updated_datasets
AseelOmer Jul 27, 2025
c96714d
added more info in the dataset readme
Alaa-Elgozouli Jul 27, 2025
ad4ad8b
added more info in the dataset readme
Alaa-Elgozouli Jul 27, 2025
84338f5
added more info in the dataset readme
Alaa-Elgozouli Jul 27, 2025
b7b78b1
updates
AseelOmer Jul 27, 2025
45e6f20
aligned datasets readme and fixed broken links in aegean_hypatia note…
Elocodes Jul 28, 2025
e11da9e
fixed formatting issues
Elocodes Jul 28, 2025
07feb4e
Merge pull request #114 from MIT-Emerging-Talent/Readme_cleanups
AseelOmer Jul 28, 2025
3816035
Data Analysis After reviewing
majdadel20 Jul 28, 2025
77223b8
Data Analysis Improved
majdadel20 Jul 28, 2025
911307e
pos
AseelOmer Jul 28, 2025
2e1d31f
Format POS tagging notebook
AseelOmer Jul 28, 2025
a60e88a
Format POS tagging notebook
AseelOmer Jul 28, 2025
e6e0aa8
Revert accidental changes to README.md
AseelOmer Jul 28, 2025
19de737
updates
AseelOmer Jul 28, 2025
9bd2d2f
updates
AseelOmer Jul 29, 2025
6134167
updates
AseelOmer Jul 29, 2025
07a20b6
updates
AseelOmer Jul 29, 2025
0ab1303
Data Analysis Readme file after reviewing multiple times
majdadel20 Jul 29, 2025
11d9543
Merge pull request #115 from MIT-Emerging-Talent/pos-tagging-nlp
Alaa-Elgozouli Jul 29, 2025
c98c0ed
Enhanced Data Analysis Readme file
majdadel20 Jul 30, 2025
2c81463
updates
AseelOmer Jul 30, 2025
4365541
Merge pull request #117 from MIT-Emerging-Talent/readme-m3
Alaa-Elgozouli Jul 31, 2025
30337f5
added more details on data preparation
Alaa-Elgozouli Jul 31, 2025
a16a09c
Fix the broken link for Data ANalysis Readme file
majdadel20 Jul 31, 2025
3dd7059
Merge pull request #103 from MIT-Emerging-Talent/analysisreadme
Alaa-Elgozouli Jul 31, 2025
c49b3e4
updates
AseelOmer Jul 31, 2025
7fb2aad
updates
AseelOmer Jul 31, 2025
ef056ce
updates
AseelOmer Aug 1, 2025
27583ef
Merge pull request #120 from MIT-Emerging-Talent/readme-m3
Alaa-Elgozouli Aug 1, 2025
0e80690
updates
AseelOmer Aug 3, 2025
af4d906
podcast script
majdadel20 Aug 10, 2025
f13eb6b
Podcast Script
majdadel20 Aug 10, 2025
b2a2eb7
podcast edit
majdadel20 Aug 10, 2025
d4854cb
Try to edinting the format
majdadel20 Aug 10, 2025
53460b7
updates
AseelOmer Aug 10, 2025
1df9365
updates
AseelOmer Aug 10, 2025
246e057
updates
AseelOmer Aug 10, 2025
36600be
Merge pull request #123 from MIT-Emerging-Talent/podcast
AseelOmer Aug 11, 2025
9e9dfd5
added communication artefact
Elocodes Aug 12, 2025
fd071b5
Merge pull request #125 from MIT-Emerging-Talent/artefact
AseelOmer Aug 12, 2025
b449a21
Add README for Communication Strategy
Rouaa93 Aug 13, 2025
f7db3c0
edit README for Communication Strategy
Rouaa93 Aug 13, 2025
3b9b6fc
Merge pull request #126 from MIT-Emerging-Talent/communication-strate…
majdadel20 Aug 13, 2025
638bc97
group represective
majdadel20 Aug 14, 2025
afe2657
Merge pull request #127 from MIT-Emerging-Talent/group
AseelOmer Aug 14, 2025
e5ba0ab
updates
AseelOmer Aug 14, 2025
0b0f698
Merge pull request #128 from MIT-Emerging-Talent/podcast
Alaa-Elgozouli Aug 14, 2025
c42e0ae
updates
AseelOmer Aug 14, 2025
258818e
Merge pull request #129 from MIT-Emerging-Talent/readme-m3
Alaa-Elgozouli Aug 15, 2025
450f9ba
updates
AseelOmer Aug 16, 2025
1a3458c
updates
AseelOmer Aug 16, 2025
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions .flake8
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
[flake8]
max-line-length = 88
extend-ignore = E203, W503, E231, E128
exclude = .git, __pycache__, .ipynb_checkpoints
27 changes: 27 additions & 0 deletions .github/workflows/lint.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,27 @@
name: Lint Python

on: [push, pull_request]

jobs:
lint:
runs-on: ubuntu-latest

steps:
- name: Checkout code
uses: actions/checkout@v3

- name: Set up Python
uses: actions/setup-python@v4
with:
python-version: '3.10'

- name: Install dependencies
run: |
python -m pip install --upgrade pip
pip install flake8 nbqa

- name: Lint .py files
run: flake8 --exclude=.venv,__pycache__ .

- name: Lint .ipynb notebooks
run: nbqa flake8 .
4 changes: 4 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -5,3 +5,7 @@ venv/
*.db
*.idea
*.ruff_cache
.env
.venv_datacleaning/
*.ipynb_checkpoints
.env
10 changes: 10 additions & 0 deletions .ls-lint.yml
Original file line number Diff line number Diff line change
@@ -1,17 +1,27 @@
version: 1

ls:
.dir: snake_case
.*: snake_case
.md: snake_case | regex:[0-9A-Z\-]+
.txt: snake_case | kebab-case
.yml: snake_case | kebab-case

rules:
filename:
regex: '^[a-z0-9_]+$'
match: true

ignore:
- .git
- .github
- .vscode
- "**/.vscode/**"
- venv
- .ruff_cache
- .pytest_cache
- __pycache__
- .ls-lint.yml
- .markdownlint.yml
- "**/node_modules"
- README.md
6 changes: 5 additions & 1 deletion .vscode/settings.json
Original file line number Diff line number Diff line change
Expand Up @@ -122,5 +122,9 @@
"source.fixAll.ruff": "explicit",
"source.organizeImports.ruff": "explicit"
}
}
},
"cSpell.words": [
"Alaa",
"stopwords"
]
}
122 changes: 122 additions & 0 deletions 0_domain_study/README.md
Original file line number Diff line number Diff line change
@@ -1 +1,123 @@
# Domain Research

## Event

In the modern time, the evolution of technology has transformed the way people
navigate their professional lives. Nowadays, online job platforms serve as a
shared environment between recruiters and job seekers.
The influence of internet has become central to recruitment strategies, hence, the
recruitment process has increasingly moved online. This shift has also brought
new challenges, as although the number of global job postings is increasing
with this shift, not all of them are legitimate.

Online job scams surged by 84% in 2023 worldwide. The rise of remote work,
online recruiting, and AI tools has made it easier for scammers to create
convincing fake job ads. This concerning trend continues to escalate, blurring
the line between legitimate opportunities and fraudulent schemes. It has become
crucial to understand what specific features in job postings and application
processes influence individuals’ engagement, and more importantly, their ability
to recognize and avoid scams.

This research aims to explore the role of visual and textual elements,
application process, along with LLMs in shaping user perceptions and actions
in the online job-seeking environment.

---

## Pattern

One emerging pattern is the use of highly attractive job descriptions and offers,
often mimicking real postings. These ads typically highlight flexible hours,
remote work, and unusually high salaries, features that immediately catch the
attention of job seekers, especially those urgently looking for income or remote
opportunities.

Simultaneously, scammers have become more sophisticated in how they design and
execute fake job listings. Leveraging LLMs,
they craft professional-sounding job descriptions and even imitate the branding
of real companies. These tools allow scammers to avoid earlier red flags such
as poor grammar or suspicious formatting, making the scams appear more credible
and trustworthy.

A major red flag in many scams involves the request for upfront payments. Whether
it's for background checks, training materials, or equipment, fraudulent
postings often ask for money early in the process with a promise of being paid back.
Patterns show that the
decision to proceed or back out often depends on how the request is framed. If the
request is made after a series of
seemingly normal interactions, people are more likely to comply, especially if
they’ve already invested time and emotional energy in the application.

---

## Structure

At a structural level, one of the key reasons online job scams are so effective
is that fraudulent processes often mimic the legitimate structure of real hiring
systems. Across industries and regions, job applications tend to follow a
standardized format. Most job seekers are familiar with being asked for personal
details such as **_full name_**, **_address, date of birth, contact information_**,
and even **_bank details_** in some cases. In many cases,
**_background checks_** and **_document verification_** are also expected steps.
These practices are so deeply rooted in professional norms that applicants rarely
question them, which makes it easier for scammers to
blend in without raising suspicion.

---

## Mental Mode

Underlying beliefs and social pressures play a significant role in sustaining
the system in which online job scams thrive. One core belief is rooted in the
reality that the global job market is increasingly competitive. With more
individuals seeking employment than there are available positions, people often
feel immense pressure to secure a job as quickly as possible. This urgency is
further amplified by the growing concerns over financial stability and societal
expectations. Therefore, job seekers may be more willing
to overlook potential red flags if it means landing a promising opportunity.

Additionally, feelings of shame and embarrassment further enforce the system.
Victims of job scams often blame themselves, fearing judgment if
they share their experience. This silence prevents others from learning about
common scam tactics and creates an environment where fraudulent practices can
continue unchecked.
Along with the fact that this issue is viewed
as individual problems, rather than systemic ones. Society often emphasizes
personal responsibility, encouraging individuals to “try harder” or
“be more careful”, instead of addressing broader structural issues.

These internalized beliefs about urgency, personal failure,
and individual responsibility, act as invisible forces that keep the scam
ecosystem alive and difficult to disrupt.

---

## Conclusion

By mapping the user journey, we can better understand how people navigate scam
encounters. When users say, “ _**What would I lose? I’ll apply and see how this
goes,**_ ” it reflects a sense of low-risk experimentation driven by desperation
or hope. Internally, they may think,
“ _**I can’t help but set high hopes incase I get the offer,**_ ”
revealing the emotional vulnerability that scammers exploit. Behaviorally, users
often proceed to apply despite some awareness of scams. Emotionally, the journey
is intense, initial excitement when
they’re contacted, followed by feelings of disappointment, panic, and
discouragement once they realize the offer was fake.

This research aims to deeply understand the
online job-seeking experience from the user's perspective, empathizing with their
motivations, behaviors, and emotional journey. By this research we're planing to
go beyond just raising awareness, and actually identify patterns to build
preventative strategies, platform-level solutions, and behavioral cues that
empower users earlier in the process.

---

## Sub Domains

- Feature engineering.
- NLP ( Natural Language Processing ).
- Human-Machine interaction.
- Machine-Machine interaction.
- LLMs.
7 changes: 7 additions & 0 deletions 0_domain_study/brainstorming.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
# Brainstorming

To access our **_group_** brainstorming:
[**Group Brainstorm**](https://docs.google.com/document/d/1GRCxNJAbXSTDNgDnKzDF3IzyhiA_VJNgygS9rDqrbrI/edit?usp=sharing).

To access our **_initial presentations_** and brainstorming:
[**The Hypatia Circle: Individual Brainstorm**](https://drive.google.com/drive/folders/1pBaj0Guxkd12CB6fYI9RAk0tbH2KZCLa?usp=sharing).
18 changes: 18 additions & 0 deletions 0_domain_study/sub_questions_and_refrences.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
# Sub-questions

- What type of features or details are included in job postings that influence
people to interact and apply?
- Given the fact that postings could be created by LLMs, how does this affect
people's ability to distinguish legitimate postings from fraud ones?
- What type of features or details are included in the application process that
influence people to continue with the process?
- What type of features or details are included in the request for upfront
payment that influence people's decision whether to pay and proceed or to
recognize this as a scam?
- If a ‘scam percentage score’ appears on the post, what is the impact of it?
Could it be helpful?

## References

- [BBC News](https://www.bbc.com/news/business-66592219).
- [User Cases](https://assets.publishing.service.gov.uk/media/61981a328fa8f5037e8ccacb/Job_Scams_-_Case_Studies.pdf).
161 changes: 161 additions & 0 deletions 1_datasets/README.md
Original file line number Diff line number Diff line change
@@ -1 +1,162 @@
# Datasets

Only one dataset was used for the analysis process of this project, that is the
**Employment Scam Aegean Dataset**. All data processing, analysis, and
observations are mainly based on this dataset, along with 500 real job posts
which were extracted from **Indeed.**

## Aegean Raw Data

This dataset can be found in **Kaggle**, referred to as **_'Recruitment Scam
Dataset'_**, or in
**EMSCAD** project **website**, referred to as **_'Employment Scam Aegean Dataset'_**.

This dataset was collected and curated by the **_Laboratory of Information and Communication
Systems Security_** at the **_University of the Aegean_** in **_Greece_**.
It contains **18,000** job postings, with around **800** labeled as fraudulent.

This dataset
was mainly designed to provide a realistic and comprehensive resource for research
on employment scams. It was collected from real online job ads between **2012**
and **2014**.
The job postings were gathered from multiple **_online_** sources, including **_job
portals_** and **_corporate websites_**.

Each entry in the dataset includes a
variety of features such as **_job title, location, department, salary range, company_**
**_profile, job description, requirements, benefits, telecommuting status, company
logo_**, **_presence, employment type, required experience, required education, industry,
function,_** and a **_class label_** indicating whether the posting is fraudulent
or not.

The dataset is publicly available and has been widely used in academic research
for developing and testing machine learning models to detect fraudulent job postings.

The main goal of this data is to clean it based on specific features that we aim
to use in the job board, organize it, refine it by Gemini, then use it to test
humans' and machines' ability to detect if the job provided is legitimate or
fraudulent when it's written by AI, since it mimics real job posts while
maintaining fraudulent posts main features.

[**File Path to Aegean Raw Data**](https://github.com/MIT-Emerging-Talent/ET6-CDSP-group-21-repo/blob/28fb2c5be79be0883c8366fb2b4bacbbec9c6809/1_datasets/aegean_raw_data)

---

## Cleaned Data

There are three files in the [`../1_datasets/cleaned_data`](https://github.com/MIT-Emerging-Talent/ET6-CDSP-group-21-repo/blob/1559fd4f70f49837b9626a46db57799e8c5a39da/1_datasets/cleaned_data)
folder:

### [Fake Job Posts](https://github.com/MIT-Emerging-Talent/ET6-CDSP-group-21-repo/blob/14894562ec2b519501aaed5b0525f54313fdfb0f/1_datasets/cleaned_data/fake_jobs.csv)

- Includes 866 fake job posts which were extracted from the Aegean raw dataset. Shape
of the dataset is (866, 11) after cleaning.

### [Real Job Posts](https://github.com/MIT-Emerging-Talent/ET6-CDSP-group-21-repo/blob/14894562ec2b519501aaed5b0525f54313fdfb0f/1_datasets/cleaned_data/real_jobs.csv)

- Includes 17014 real job posts which were extracted from the Aegean raw dataset.
Shape of the dataset is (17014, 11) after cleaning.

### [LLM-Refined Fake Job Posts](https://github.com/MIT-Emerging-Talent/ET6-CDSP-group-21-repo/blob/1559fd4f70f49837b9626a46db57799e8c5a39da/1_datasets/cleaned_data/llm_refined_fake_posts2.csv)

- Includes 866 fake job posts which were all LLM-refined. Shape of the dataset
is (866, 15) after cleaning.

---

## Aegean500_Hypatia500 Datasets Folder

In addition to the primary analysis focusing on the Aegean dataset, a
complementary analyses was done to explore the detection capabilities of four
traditional Machine Learning Models when given fake jobs refined by AI and real
jobs posted in the AI era.

### A note on the datasets that were used for the analyses

There are three files in the [`../1_datasets/aegean500_vs_hypatia500_datasets`](https://github.com/MIT-Emerging-Talent/ET6-CDSP-group-21-repo/blob/e46c53bf17c3d608c8e67b607300d9faf4b6043e/1_datasets/aegean500_vs_hypatia500_datasets)
folder:

### [**Fake Job Posts**](https://github.com/MIT-Emerging-Talent/ET6-CDSP-group-21-repo/blob/e46c53bf17c3d608c8e67b607300d9faf4b6043e/1_datasets/aegean500_vs_hypatia500_datasets/aegean500_fakejobs.csv)

- The (866, 17) fake jobs in the Aegean dataset contains lots of missing values
and needed to be cleaned for the analysis. A research of modern day job posts
structure revealed the following as core features - **job_id, job_title, location,
job_description, benefits** (with benefits often lumped into the job description).

- These features were retained and used to remove missing values from the dataset,
achieving a (500, 6) shape of fake jobs.

- Hence the difference between this file and the one in [`../1_datasets/cleaned_data/fake_jobs.csv`](https://github.com/MIT-Emerging-Talent/ET6-CDSP-group-21-repo/blob/14894562ec2b519501aaed5b0525f54313fdfb0f/1_datasets/cleaned_data/fake_jobs.csv)
is that they both followed different data cleaning processes.

### [**Real Job Posts**](https://github.com/MIT-Emerging-Talent/ET6-CDSP-group-21-repo/blob/e46c53bf17c3d608c8e67b607300d9faf4b6043e/1_datasets/aegean500_vs_hypatia500_datasets/hypatia500_realjobs.csv)

- contains 500 real jobs scraped from the job board **Indeed** (_Jobs were
sorted by 'recently posted' on the board, and retrieved between 7/19/2025 - 7/20/2025_).
The chrome extention "webscraper.io" was utilized for the scrapping and these
key features - **Job_title, location, job description and job link** were
scrapped, with source-date retained as metadata.

- What informed the 'type' of real jobs to scrape?
- The fake jobs word cloud showed that certain job titles - Data Entry,
Engineer, Customer service, and Entry clerk, were dominant. These became the
key jobs that were searched and scrapped from Indeed.

- Hence, the difference between this file and the one in [`../1_datasets/cleaned_data/real_jobs.csv`](https://github.com/MIT-Emerging-Talent/ET6-CDSP-group-21-repo/blob/14894562ec2b519501aaed5b0525f54313fdfb0f/1_datasets/cleaned_data/real_jobs.csv)
is the fact that they are real jobs of different times,
(aegean (2012 - 2014), hypatia (2025))

### [**LLM-Refined Fake Job Posts**](https://github.com/MIT-Emerging-Talent/ET6-CDSP-group-21-repo/blob/e46c53bf17c3d608c8e67b607300d9faf4b6043e/1_datasets/aegean500_vs_hypatia500_datasets/aegean500_fakejobs_llmrefined.csv)

- The 500 job descriptions were fed to Gemini 2.5 flash with a robust prompt
aimed at modernizing the fake jobs. Specifically, this **"Add appealing but
potentially exaggerated benefits/responsibilities."** was part of the prompt to
bring the job descriptions up to modern standard.

- The refinement task was split between team members in this file:
- [`/1_datasets/fakejobs_to_refine`](https://github.com/MIT-Emerging-Talent/ET6-CDSP-group-21-repo/blob/08990371387dcddd06fb6f3361478bf4c33d45fb/1_datasets/fakejobs_to_refine).
- The refined version is seen here:
- [`../1_datasets/fakejobs_refined`](https://github.com/MIT-Emerging-Talent/ET6-CDSP-group-21-repo/blob/8251dfa7db2ae2e0b35c3e619dd3c7f6e52af037/1_datasets/fakejobs_refined).
- The 6 refined batches were recombined to get the aegean500_fakejobs_llmrefined.csv
- Hence, the difference between this file and [`../1_datasets/cleaned_data/llm_refined_fake_posts2.csv`](https://github.com/MIT-Emerging-Talent/ET6-CDSP-group-21-repo/blob/1559fd4f70f49837b9626a46db57799e8c5a39da/1_datasets/cleaned_data/llm_refined_fake_posts2.csv)
is that they both read and refined files that followed different cleaning processes.

### All scripts related to this complementary analyses

- [fake jobs extraction script](https://github.com/MIT-Emerging-Talent/ET6-CDSP-group-21-repo/blob/main/2_data_preparation/fake_jobs_extraction_script.ipynb)
- [fake jobs refinement script](https://github.com/MIT-Emerging-Talent/ET6-CDSP-group-21-repo/blob/main/2_data_preparation/fake_jobs_ai_refinement_script.ipynb)
- [batch recombination script](https://github.com/MIT-Emerging-Talent/ET6-CDSP-group-21-repo/blob/main/2_data_preparation/recombine_aegean500_batches.py)
- [data exploration script](https://github.com/MIT-Emerging-Talent/ET6-CDSP-group-21-repo/blob/main/3_data_exploration/aegean_hypatia_datasets_exploration.ipynb)
- [data analysis script](https://github.com/MIT-Emerging-Talent/ET6-CDSP-group-21-repo/blob/main/4_data_analysis/aegean_hypatia_datasets_analysis.ipynb)

---

## Past Relevant Studies

The research, **_Assessing AI vs Human-Authored Spear Phishing SMS Attacks: An
Empirical Study_**, a **2025** study, was conducted by researchers at **_Brigham
Young University._**

To collect data, participants were recruited and asked to provide personal
information that could be used to personalize phishing messages.

Meanwhile, on the other hand, both human and AI authors used this personal information
to craft spear phishing SMS messages tailored to each participant.

There were mainly **_four_** steps for the data collection procedure. **_Firstly_**
participants were shown a set of personalized SMS messages (some human-authored,
some AI-generated).
**_Secondly_**, they were asked to rank the messages by how convincing they found
them. **_Thirdly_**, participants also provided qualitative feedback on each message.
**_Lastly_**, they were then asked to guess whether each message was written by a
human or by AI.

This experiment measured how convincing the messages were (both written by humans
and AI) and whether participants could distinguish between human and AI authorship.

There was **_no study_** that specifically tested humans' ability to differentiate
if a job post is legitimate or fraudulent when the post is AI-generated that we
know of, hence we're planning to use this study as a foundation for the process
of building and analyzing the job board.

[**Past Relevant Studies File Path.**](https://github.com/MIT-Emerging-Talent/ET6-CDSP-group-21-repo/blob/549942e9039edafdae73dff7d904ec97fa432148/1_datasets/past_relevant_studies.md)
Loading
Loading