The Smart Expense Categorizer is a Streamlit-based web app that automatically categorizes bank transactions from uploaded CSV or Excel files. It uses a machine learning model (TF-IDF and Naive Bayes) to analyze the transaction description and predict the expense category. Users can upload their bank statement, view the processed data, analyze spending patterns, and download a categorized report. The app helps simplify personal expense tracking by automating the process of organizing transactions.
flowchart TB
subgraph PresentationLayer [Presentation Layer]
ST[Streamlit UI]
end
subgraph BusinessLayer [Business Layer]
ModelTrainer[Model Trainer]
Categorizer[Expense Categorizer]
ArrowCompat[Arrow Compatibility Handler]
end
subgraph DataAccessLayer [Data Access Layer]
FileUpload[File Upload/Parsing]
DataFrame[Bank Statement DataFrame]
end
FileUpload --> DataFrame
DataFrame --> ArrowCompat
ArrowCompat --> Categorizer
ModelTrainer --> Categorizer
Categorizer --> ST
FileUpload --> ST
-
Purpose: Provides the interactive web interface for file upload, column selection, categorization, analytics, and download.
-
Key Properties:
- Title and description
- File uploader widget
- Data preview table
- Column selection dropdowns
- Categorize button
- Analytics visualization (bar chart, value counts)
- Download button for results
- Error/info dialogs
-
Key Methods:
st.file_uploader(): Handles file input.st.selectbox(): Lets users pick description/amount columns.st.button("Categorize Expenses"): Triggers categorization.st.dataframe(),st.bar_chart(),st.download_button(): For data display and interaction.
- Purpose: Prepares and fits the machine learning pipeline on startup using labeled example data.
- Key Methods:
train_model(): Trains a Scikit-learn Pipeline (TF-IDF + MultinomialNB) on Indian-context phrases and categories.
- Purpose: Applies the trained model to user-uploaded transaction descriptions and assigns categories.
- Key Methods:
model.predict(): Predicts a category for each transaction.
- Purpose: Ensures the Pandas DataFrame can be serialized for Streamlit/Arrow, preventing serialization errors.
- Key Methods:
make_arrow_compatible(frame: pd.DataFrame): Converts problematic columns to string or appropriate types.
- Purpose: Reads and parses user-uploaded CSV or Excel files into Pandas DataFrames.
- Key Methods:
- Conditional logic for handling
.csvand.xlsxfiles, including encoding fallback.
- Conditional logic for handling
- Purpose: Holds the uploaded data, intermediate, and processed data throughout the session.
- Properties:
| Property | Type | Description |
|---|---|---|
training_data |
List[Tuple[str, str]] |
List of (description, category) pairs for model training. |
- Properties: User-supplied columns (e.g., "Date", "Description", "Amount") plus "Predicted Category" added after processing.
No HTTP API endpoints are defined in this code.
All functionality operates locally in the Streamlit session.
sequenceDiagram
participant U as User
participant ST as Streamlit UI
participant File as File Upload Handler
participant Trainer as Model Trainer
participant Cat as Categorizer
participant Arrow as Arrow Handler
U->>ST: Open app in browser
U->>ST: Upload CSV/Excel
ST->>File: Parse and read file
File-->>ST: DataFrame preview
U->>ST: Select description/amount columns
U->>ST: Click "Categorize Expenses"
ST->>Cat: Pass DataFrame with selected columns
Cat->>Trainer: Use trained model
Cat->>Arrow: Ensure Arrow compatibility
Arrow-->>Cat: Arrow-safe DataFrame
Cat-->>ST: Categorized DataFrame
ST-->>U: Show results, analytics, download option
- Initial: Waiting for file upload.
- File Uploaded: Data previewed, columns selectable.
- Categorization Running: Spinner shown, disables UI.
- Result: Display of categorized data and analytics.
- Error: Error/information box shown.
- Pandas: For data manipulation and file reading.
- Scikit-learn: For NLP feature extraction and classification.
- Openpyxl: For Excel file processing.
- Streamlit: For UI and app orchestration.
- Screen Views: Main dashboard, data preview, analytics section.
- User Actions:
- File upload
- Column selection
- Categorization trigger
- CSV download
| Class / Function | Location | Responsibility |
|---|---|---|
train_model |
expense_app.py | Trains the TF-IDF + Naive Bayes pipeline |
make_arrow_compatible |
expense_app.py | Ensures DataFrame serialization compatibility |
| Streamlit UI Components | expense_app.py | Orchestrate user workflow |
- File Type Detection: Only
.csvand.xlsxfiles accepted; others trigger error. - Encoding Fallback: Attempts alternate encoding for CSVs if default fails.
- Arrow Serialization Protection: Converts problematic columns to string when Arrow cannot serialize.
- Column Existence Checks: Ensures selected columns exist; else, stops with error.
- Amount Parsing: Handles currency signs, commas, and conversion errors gracefully.
- General Exception Handling: Catches and displays any error during file processing.
Example:
if file_extension == ".csv":
try:
df = pd.read_csv(uploaded_file)
except UnicodeDecodeError:
df = pd.read_csv(uploaded_file, encoding="latin1")
elif file_extension == ".xlsx":
df = pd.read_excel(uploaded_file)
else:
st.error("Unsupported file format!")
st.stop()- No explicit caching is implemented in the code.
- Model is trained on every script run; data is processed per session.
- streamlit
- pandas
- scikit-learn
- openpyxl
- joblib
- Test all file upload scenarios: CSV, XLSX, invalid files, encoding variants.
- Try missing/extra columns: Ensure errors are handled.
- Validate categorization: Known descriptions yield correct categories.
- Check download: Downloaded file matches processed data.
The application does not expose HTTP API endpoints; all logic is contained in the Streamlit session.
- π Upload CSV or Excel bank statements
- π€ Automatic AI-powered categorization (trained on Indian merchant/expense patterns)
- π§ NLP engine: TF-IDF + Naive Bayes for smart predictions
- π Spending analytics: Bar charts & category breakdowns
- π₯ Download categorized results
- π‘οΈ Robust error handling and compatibility with complex input files
- βοΈ Configurable column selection
- π‘ Sample statement and clear guidance
- Python 3.11
- Streamlit (UI and app framework)
- pandas (data handling)
- scikit-learn (NLP and classification)
- openpyxl (excel file processing)
- joblib (model handling)
-
Install dependencies:
pip install streamlit pandas scikit-learn openpyxl
-
Run the app:
python -m streamlit run expense_app.py
-
Use the web UI:
- Upload your bank statement file (
.csvor.xlsx) - Select the transaction description and (optionally) amount columns
- Click Categorize Expenses
- Explore your spending analytics and download the results as CSV
- Upload your bank statement file (
-
Clone the repository:
git clone https://github.com/yourusername/smart-expense-categorizer.git cd smart-expense-categorizer -
Install requirements:
pip install -r requirements.txt
-
Launch the app:
python -m streamlit run expense_app.py
-
(Optional) Deploy to Streamlit Cloud:
- Push your code to GitHub.
- Go to streamlit.io/cloud and connect your repository.
- Add support for PDF and bank-specific statement formats
- Expand training data for more granular categories
- Persist user models for personalized learning
- Add user authentication for multi-user support
- Visualize transactions on a timeline or calendar view
This project is licensed under the MIT License.