Project was built in 48 hours for Vietnam AI Innovation Challenge Hackathon
National Innovation Center (NIC) • Innovation Track
An AI assistant that helps startups, FDI enterprises, and high-tech companies discover relevant government policies, incentives, and funding opportunities by:
- Searching and interpreting applicable laws, decrees, circulars, and government support programs based on the company’s needs.
- Continuously monitoring policy updates and newly announced funding programs.
- Assisting organizations in preparing and completing grant and funding application documents.
P2B là một không gian làm việc số (workspace) tích hợp trí tuệ nhân tạo (AI-native) toàn diện, giúp doanh nghiệp tự động hóa quy trình phân tích tài liệu pháp lý, đối chiếu tiêu chuẩn tài trợ và chuẩn bị hồ sơ ứng tuyển chất lượng cao.
Lập hồ sơ doanh nghiệp có trích dẫn nguồn, đối chiếu tự động với hàng trăm chính sách nhà nước và xuất đơn ứng tuyển chuẩn chỉ chỉ trong vài phút.
Trải nghiệm phiên bản thử nghiệm trực tuyến tại: Bản Demo không còn nữa!
- Tính năng cốt lõi & Quy trình Người dùng
- Chi tiết Kiến trúc Hệ thống Vector (Vector Engine Architecture)
- Tổng quan Công nghệ (Tech Stack)
- Hướng dẫn Thiết lập Local (Full-Stack)
- Cấu hình Biến môi trường
- Tải tài liệu: Doanh nghiệp tải lên các văn bản chứng minh pháp lý (Giấy phép kinh doanh, Báo cáo tài chính, Chứng nhận đầu tư...) lên vùng lưu trữ bảo mật (Supabase Private Storage).
- Trích xuất cấu trúc bằng AI: Hệ thống sử dụng Microsoft MarkItDown để chuyển đổi PDF/Docx thành Markdown, sau đó gọi mô hình
gemini-3.1-flash-liteđể trích xuất các thông tin định dạng sẵn. - Đảm bảo không ảo giác (Zero Hallucination): Mỗi thông tin trích xuất bắt buộc phải có câu trích dẫn (
quote) khớp chính xác từng ký tự trong văn bản gốc mới được đưa vào trạng thái Chờ Duyệt (NEEDS_REVIEW). - Duyệt và Cập nhật: Người dùng đối chiếu nguồn, xác nhận thông tin đúng để tạo phiên bản Hồ sơ Doanh nghiệp chính thức mới.
- Không phụ thuộc vào AI (AI-free Evaluation): Động cơ sử dụng mã nguồn Go để tính toán chính xác điều kiện của doanh nghiệp so với tiêu chuẩn chính sách (gồm các so sánh chuỗi, ngày tháng, số học phức tạp như
EQ,IN,CONTAINS,GT,GTE,LT,LTE,EXISTS,DATE_BEFORE,DATE_AFTER). - Phân loại trạng thái: Kết quả trả về gồm
MET(Đạt),NOT_MET(Không đạt) hoặcMISSING_INFO(Thiếu thông tin chứng minh). - Xếp hạng tự động: Gợi ý các chính sách hỗ trợ/tài trợ có tỷ lệ khớp cao nhất cho doanh nghiệp.
- Multi-business Tenants: Một tài khoản có thể sở hữu nhiều business workspace; mỗi request production phải qua membership check để bảo mật dữ liệu. Nút
Cập nhật tài liệutạo refresh job chỉ phân tích PDF mới, bảo toàn Passport hiện tại. - Layout-preserving PDF extraction: Văn bản PDF chứa bảng biểu được bổ sung công cụ
pdftotext -layoutgiữ quan hệ nhãn-giá trị; PDF scan/ít text chạy OCR fallback bằngpdftoppm+tesseract(OCR_LANGUAGESmặc địnhvie+eng). - Completeness Pass & Quality Gates: Chạy completeness pass có mục tiêu cho các field còn thiếu và ghi nhận logs chất lượng (raw/valid/rejected candidates) thay vì logs các trường dữ liệu nhạy cảm.
- Opportunity Matching Cache: Kết quả phân tích đối chiếu tiêu chuẩn tài trợ được lưu trữ lưu động trong bảng
match_runscủa PostgreSQL. Cache tự động xóa và tính toán lại khi người dùng thay đổi dữ liệu trong Passport, giúp giảm thiểu thời gian chờ đợi. - Watchlist Toggles: Cho phép doanh nghiệp tùy biến bộ lọc cảnh báo cho từng workspace với 4 danh mục: Chính sách mới, Thay đổi thời hạn, Chứng cứ hết hạn (Stale), và Hồ sơ cận hạn. Trạng thái cấu hình được lưu trực tiếp dưới dạng JSONB.
- Automated Crawler Scheduler: Bộ quét chạy ngầm định kỳ mỗi 1 giờ (và chạy ngay khi khởi động hệ thống) để đối chiếu thay đổi của văn bản pháp luật:
- Phát hiện chính sách mới phù hợp hồ sơ doanh nghiệp.
- Phát hiện các cập nhật thay đổi ngày hạn chót của chính sách hiện tại.
- Đối chiếu ngày hết hiệu lực của các văn bản pháp luật với mã băm nội dung (
content_hash) của tài liệu minh chứng trong Passport để tự động cảnh báo chứng cứ cũ/hết hiệu lực.
- Mẫu đơn tùy biến: Hỗ trợ tải lên và phân tích các mẫu đơn PDF/Docx của chính phủ, tự động trích xuất các trường thông tin cần điền bằng Gemini.
- Application Draft Cache: Toàn bộ bản nháp và tiến trình chuẩn bị đơn được lưu trữ động trong bảng
application_draft_cachecủa PostgreSQL, đảm bảo dữ liệu không bị mất khi tải lại trang hoặc đổi workspace.
- Sinh Checklist tự động: Dựa trên các điều kiện luật của chính sách, hệ thống tạo checklist các văn bản bắt buộc cần nộp.
- Khóa hồ sơ theo bằng chứng: Chỉ khi toàn bộ các trường thông tin quan trọng được xác minh nguồn dẫn, đơn ứng tuyển mới cho phép hoàn tất và phê duyệt.
- Xuất đơn PDF: Tự động tạo và điền các mẫu đơn ứng tuyển dạng PDF/Docx chứa đầy đủ chữ ký số và bằng chứng đi kèm để nộp lên cơ quan quản lý.
Hệ thống Vector của P2B được thiết kế chuyên biệt để xử lý dữ liệu lớn luật pháp Việt Nam (VBPL) với hiệu năng cao nhất trên GPU local và tài nguyên tiết kiệm nhất trên CPU production.
- Bộ lọc Whitelist: Tự động lọc bộ dữ liệu tmquan/vbpl-vn trên Hugging Face (dựa trên cơ sở dữ liệu pháp luật chính thức vbpl.vn) theo cơ quan ban hành liên quan đến kinh tế/doanh nghiệp (Bộ Tài chính, Bộ Kế hoạch và Đầu tư, Ngân hàng Nhà nước...) kết hợp keyword pháp lý hỗ trợ.
- Phục hồi văn bản bị thiếu: Với các văn bản thiếu body text, pipeline tự động kết nối qua cổng thông tin MoJ API gateway (
https://vbpl-bientap-gateway.moj.gov.vnliên kết với vbpl.vn) để tải bản thảo XML đầy đủ. - Article-level Chunking: Tách văn bản luật chính xác theo từng Điều (
Điều X), tự động xử lý inline headings lỗi không xuống dòng. Trích xuất được 19,488 chunks từ 682 văn bản luật cốt lõi, chiều dài trung bình 1,769 ký tự/chunk. - Mã hóa đa luồng: Chạy SentenceTransformer
intfloat/multilingual-e5-basesong song thông quaThreadPoolExecutorvàThreadedConnectionPooltrên GPU NVIDIA (Có hỗ trợ CUDA) tại local với khóa GPU lock an toàn.
Khi chạy thực tế trên môi trường CPU của Railway (không có GPU và giới hạn RAM):
- Kiến trúc CGO-free: Go backend giao tiếp với script trợ lý Python (
calculate_embeddings.py) thông qua luồng Standard Input (stdin) để tránh quá giới hạn ký tự dòng lệnh. - ONNX Runtime & Tokenizers: Thay thế toàn bộ PyTorch và Transformers bằng
onnxruntimevà Rust-basedtokenizersđể tối ưu hóa bộ nhớ: giảm thiểu tài nguyên tiêu thụ từ 1.5GB RAM xuống <60MB RAM. - 8-bit Quantization: Sử dụng phiên bản mô hình E5 đã lượng hóa 8-bit (
model_quantized.onnx) giúp giảm dung lượng ổ đĩa từ 278MB xuống 141MB và tăng gấp đôi tốc độ suy luận CPU với độ tương quan chính xác đạt 99.9%. - Preload model khi build image: Docker build tải model & tokenizer, chạy một inference kiểm tra, rồi đóng cache vào image. Request production đầu tiên không phụ thuộc mạng hoặc thời gian tải model.
- Chỉ mục vector HNSW (
vector_cosine_ops) được tối ưu hóa bằng cách tắt song song hóa bảo trì (SET max_parallel_maintenance_workers = 0) để xây dựng chỉ mục tuần tự, bypass thành công giới hạn bộ nhớ chia sẻ (/dev/shm) của các container ảo hóa trên Railway. - Policy matching dùng Reciprocal Rank Fusion giữa PostgreSQL full-text rank và cosine similarity 768 chiều, sau đó hợp nhất với kết quả rule engine đã review.
- Backend: Go 1.26+, go-chi/chi router,
pgx/v5PostgreSQL client, embedded SQL migrations. - Frontend: React 19, TypeScript, Vite, TanStack Query v5, Motion (animations), Radix UI, Vanilla CSS (biến CSS tùy chỉnh).
- Dịch vụ ngoài & Lưu trữ: Supabase (Xác thực, Storage riêng tư), Railway PostgreSQL 17 (pgvector 0.8.2).
- Trích xuất & AI: Microsoft MarkItDown CLI, Google Gemini API (
gemini-3.1-flash-lite). - Môi trường chạy sản phẩm (Production): Docker (chứa Poppler, Tesseract
vie+eng, LibreOffice, ClamAV) chạy trên Railway; SPA hosting chạy trên Vercel.
- Go 1.24+ (hoặc 1.26)
- Node.js 22+ & npm
- Python 3.10+ (cần thiết cho local extraction worker chạy thư viện
markitdown) - Cài đặt công cụ CLI và các thư viện cần thiết:
pip install markitdown[pdf]==0.1.6 onnxruntime tokenizers numpy
Trong chế độ phát triển mặc định (DEV_AUTH=true), hệ thống sử dụng bộ lưu trữ bộ nhớ (in-memory) giả lập. Bạn không cần kết nối PostgreSQL hay Supabase để chạy ứng dụng!
- Sao chép cấu hình môi trường:
cp .env.example .env
- Khởi động Backend (Go API):
make api # Hoặc: cd api && DEV_AUTH=true go run ./cmd/api - Khởi động Frontend (React/Vite):
Mở một terminal mới và chạy:
make web # Hoặc: cd web && npm install && npm run dev - Truy cập ứng dụng tại địa chỉ:
http://localhost:5173. Header mặc địnhX-Workspace-IDđược dùng để phân chia workspace độc lập.
- Chạy toàn bộ unit test:
make test - Kiểm tra lỗi cú pháp và lint:
make lint
- Biên dịch dự án:
make build
Nếu muốn chạy với PostgreSQL cục bộ hoặc môi trường Supabase:
- Đảm bảo đặt
DEV_AUTH=falsevàVITE_DEV_AUTH=falsetrong file.env. - Điền đầy đủ thông tin kết nối
DATABASE_URL(PostgreSQL), các khóaSUPABASE_*vàGEMINI_API_KEY. - Chạy lệnh migrate để tạo bảng:
make migrate
Chi tiết các biến môi trường cấu hình tại file .env:
DEV_AUTH: Thiết lậptrueđể bỏ qua xác thực Supabase JWT, hữu dụng cho phát triển local.DATABASE_URL: Đường dẫn kết nối PostgreSQL của Railway (chỉ dùng khiDEV_AUTH=false).SUPABASE_URL/SUPABASE_SECRET_KEY: Dùng cấu hình xác thực và vùng chứa Storage.GEMINI_API_KEY: API Key kết nối dịch vụ Google Gemini.P2B_MODEL_CACHE_DIR: Thư mục lưu trữ cache model ONNX (mặc định sẽ lưu tại~/.cache/p2b-embeddings).
Dự án sử dụng cơ sở dữ liệu pháp luật Việt Nam từ các nguồn dữ liệu công khai:
- Dữ liệu chính thức từ cổng thông tin: vbpl.vn (Cổng thông tin Pháp điển / Bộ Tư pháp).
- Bộ dữ liệu Hugging Face: tmquan/vbpl-vn của tác giả
tmquan.
P2B is an AI-native workspace designed to help companies streamline the grant application process by automating document extraction, checking eligibility criteria through a deterministic rule engine, and preparing application checklists.
Instantly build verified Company Passports, automatically evaluate matching grants, and generate audit-ready applications with clear evidence provenance in minutes.
- Core Features & User Flows
- Vector Engine Architecture
- Tech Stack Overview
- Full-Stack Local Development Setup
- Environment Configuration
- Document Upload: Users securely upload business evidence (Business Licenses, Financial Statements, Investment Certificates) straight to Supabase Private Storage.
- AI-Structured Extraction: The pipeline uses Microsoft MarkItDown to parse the documents into clean Markdown, then calls Gemini 3.1 Flash-Lite to extract canonical fields.
- Zero Hallucinations: Every candidate field must map to an exact quote found in the source text to prevent AI hallucinations. Unmatched/unverified facts are flagged as
NEEDS_REVIEW. - Verification Workflow: Users review the facts, resolve conflicts, and promote them to create a versioned, immutable Company Passport.
- AI-Free Matching: Go-based engine evaluates fields against policies using precise comparison operators (like
EQ,IN,CONTAINS,GT,GTE,LT,LTE,EXISTS,DATE_BEFORE,DATE_AFTER). - Evaluation Statuses: Tracks status as
MET,NOT_MET, orMISSING_INFO(if a field is unconfirmed). - Smart Ranking: Grants are sorted dynamically based on matching criteria scores.
- Multi-business Tenants: Owners can manage and switch between multiple business workspaces with strict tenant isolation. Document update triggers incremental PDF refresh jobs without destroying existing facts.
- Layout-preserving PDF extraction: Combines layout text output for tables (
pdftotext -layout) with custom OCR fallback usingpdftoppm+tesseract(OCR_LANGUAGESdefaults tovie+eng). - Completeness Pass & Quality Gates: Target checks resolve missing facts and audit quality (raw/valid/rejected counts) without logging sensitive quotes.
- Opportunity Matching Cache: Matching scores and criteria results are cached persistently inside PostgreSQL
match_runstable, avoiding redundant runs. The cache automatically invalidates and refreshes whenever the Company Passport is updated. - Interactive Watchlist: Strict workspace preferences for alerts can be toggled interactively for four distinct categories (New Policies, Deadline Changes, Stale Evidence, and Upcoming Deadlines), stored directly as a JSONB preferences column.
- Scheduled Crawler Worker: A background scheduler process executes a rules pass on startup and every 1 hour to:
- Identify new policies matching the workspace's support profile.
- Detect deadline alterations on active policies.
- Cross-reference expired legal document versions with active Passport evidence content hashes to flag stale evidence.
- Custom Form Templates: Users can upload custom government PDF/Docx form templates, automatically parsing and extracting template placeholders using Gemini structure extraction.
- Application Draft Cache: Draft sections and progress state are persisted dynamically inside
application_draft_cachein PostgreSQL to protect progress across browser reloads or workspace swaps.
- Automated Checklist: Generates document and info check-items mapped to policy requirements.
- Gated Actions: Ensures application submission is locked until necessary evidence has been reviewed and confirmed.
- PDF Package Export: Compiles confirmed fields and checklists to export a structured PDF ready for submission.
The P2B Vector Engine is engineered to parse the Vietnamese legal document corpus (VBPL) efficiently using local GPU for population and resource-restricted CPU runtime for production.
- Whitelist Filtering: Filters tmquan/vbpl-vn Hugging Face dataset (originating from vbpl.vn) by economic ministries and support keywords.
- Empty Text Resolution: Connects to the official MoJ gateway (
https://vbpl-bientap-gateway.moj.gov.vnlinked to vbpl.vn) to scrape full XML drafts for empty records. - Article-level Chunking: Splits documents on semantic article levels (
Điều X), cleaning up inline headings. Produced 19,488 chunks from 682 documents with a mean length of 1,769 characters. - Multithreaded Encoding: Concurrently embeds chunks using
ThreadPoolExecutor,ThreadedConnectionPoolandintfloat/multilingual-e5-baseon local NVIDIA RTX GPUs (With CUDA Support) with thread-safe locking.
- CGO-Free Subprocess Wrapper: Go executes
calculate_embeddings.pyusing standard inputs (stdin) to bypass command-line size boundaries. - ONNX Runtime & Tokenizers: Swaps PyTorch/Transformers with optimized C++/Rust implementations, slashing RAM consumption from 1.5GB to <60MB.
- 8-bit Quantized Model: Loads
model_quantized.onnxto cut storage footprint in half (141MB vs 278MB) and double inference speed on CPU with 99.9% cosine similarity preservation. - Build-Time Model Preload: Docker downloads the model and tokenizer and executes a verification inference during image build. The first production request has no model-download dependency.
- The pgvector HNSW cosine index was successfully generated by forcing sequential builds (
SET max_parallel_maintenance_workers = 0), preventing shared memory segment allocation failures under Railway container constraints. - Policy matching applies Reciprocal Rank Fusion to PostgreSQL full-text rank and 768-dimensional cosine similarity, then merges retrieved documents with reviewed rule-engine results.
- Backend: Go 1.26+, go-chi/chi router,
pgx/v5PostgreSQL client, embedded SQL migrations. - Frontend: React 19, TypeScript, Vite, TanStack Query v5, Motion (animations), Radix UI, Vanilla CSS (using CSS variables).
- External Services & Databases: Supabase (Auth, Private Storage), Railway PostgreSQL 17 (pgvector 0.8.2).
- Extraction & AI: Microsoft MarkItDown CLI, Google Gemini API (
gemini-3.1-flash-lite). - Production Hosting: Docker-based containers (containing Poppler, Tesseract
vie+eng, LibreOffice, ClamAV) deployed to Railway; Vercel for SPA web hosting.
- Go 1.24+ (or 1.26)
- Node.js 22+ & npm
- Python 3.10+ (required for local extraction worker running
markitdown) - Install python libraries:
pip install markitdown[pdf]==0.1.6 onnxruntime tokenizers numpy
By default, the workspace is configured to use development mode (DEV_AUTH=true) which runs with in-memory database adapters. You do not need a live PostgreSQL database or Supabase setup to run and play with the app locally!
- Clone the environment file:
cp .env.example .env
- Start the Go API Server:
make api # Or: cd api && DEV_AUTH=true go run ./cmd/api - Start the React Web Client:
In a new terminal window, run:
make web # Or: cd web && npm install && npm run dev - Open
http://localhost:5173in your browser. The defaultX-Workspace-IDheader is used to isolate workspaces.
- Run all tests:
make test - Run syntax and lint checks:
make lint
- Build the full stack:
make build
If you want to run with local PostgreSQL or live Supabase services:
- Set
DEV_AUTH=falseandVITE_DEV_AUTH=falsein your.envfile. - Supply real credentials for
DATABASE_URL(PostgreSQL),SUPABASE_*credentials, andGEMINI_API_KEY. - Apply migrations to initialize the database schema:
make migrate
Key environment configurations inside .env:
DEV_AUTH: Set totrueto bypass Supabase JWT validation, ideal for quick local development.DATABASE_URL: Connection string for Railway PostgreSQL instance (used whenDEV_AUTH=false).SUPABASE_URL/SUPABASE_SECRET_KEY: Used for authentication and bucket signing operations.GEMINI_API_KEY: API Key for Google Gemini services.P2B_MODEL_CACHE_DIR: Local persistent cache folder path for model files (defaults to~/.cache/p2b-embeddings).
This project uses Vietnamese legal documents from open public sources:
- Official data from: vbpl.vn (Ministry of Justice of Vietnam).
- Hugging Face dataset: tmquan/vbpl-vn by
tmquan.

