GovInsight-AI 是一个基于 大语言模型 (LLM) 的政务热线工单质量检测系统。它旨在解决政务热线(如 12345)中“通话录音”与“话务员录入工单”一致性校验的痛点。
传统的人工质检效率低、标准不一,且难以发现隐蔽的语义篡改。GovInsight-AI 通过自动比对录音转写与工单记录,精准识别关键信息缺失、语义偏差和风险降级等问题,并提供智能化的修正建议,大幅提升质检效率与准确性。
在政务服务热线(如 12345)的日常运营中,工单记录质量直接关系到群众诉求的办理效率和满意度。然而,传统的人工质检模式面临着巨大挑战:
- ⚡️ 效率低下:海量的通话录音和工单记录,人工抽检率通常不足 5%,大量问题工单成为“漏网之鱼”。
- 📏 标准不一:不同质检员的主观判断差异大,难以形成统一、公正的评价体系。
- 🙈 隐蔽篡改:话务员为了规避考核,可能将“投诉”私自改为“咨询”,或故意漏记群众的激进言辞,人工难以逐一核对录音。
- 📉 反馈滞后:质检通常是事后进行(T+1甚至T+7),无法在工单流转前及时拦截和修正。
GovInsight-AI 正是为解决上述痛点而生,它将 LLM 的语义理解能力引入质检环节,实现全量、实时、客观的智能检测。
GovInsight-AI 不仅仅是一个打分工具,更是一个智能辅助助手。
系统基于以下四个核心维度对工单进行深度扫描:
- 完整性 (Completeness):检测是否遗漏时间、地点、涉事对象、具体诉求等关键要素。
- 一致性 (Consistency):(核心能力) 比对录音与工单,发现语义篡改、事实偏差或性质变更(如“投诉”变“咨询”)。
- 处理规范性 (Handling Type):(新增) 智能研判诉求性质,精准区分“直接办结”与“转办”。例如,需要线下核实的投诉严禁直接办结,必须转派职能部门。
- 规范性 (Clarity):评估表述是否清晰、专业,是否存在语病、歧义或口语化表达。
- 风险敏感性 (Risk Awareness):识别是否忽视了群众的激烈情绪、重复投诉历史或潜在的舆情升级风险。
拒绝“黑盒”评判!系统会展示 AI 的完整推理过程(Chain of Thought):
"用户在录音中明确提到了‘已经是第三次投诉了’,但工单描述中未记录此信息,这属于关键信息遗漏,且降低了问题的紧迫性..." 这种可解释性让质检员能够快速复核并信任 AI 的判断。
引入置信度 (Confidence) 机制,将工单分为三类:
- ✅ 自动采信 (Auto-Pass):置信度 ≥ 0.85 且无风险的工单,直接通过,无需人工介入。
- 👀 抽检复核 (Sampling):置信度在 0.70 - 0.84 之间的工单,进入抽检池。
- 🚨 强制复核 (Mandatory Review):置信度 < 0.70 或存在高风险(如情绪激进)的工单,强制要求人工复核。
当发现质量问题时,AI 不仅会报错,还会自动重写一份标准的工单。 系统提供直观的 Diff 视图,高亮显示原工单与 AI 建议工单的差异,话务员或质检员可一键采纳建议。
(此处建议插入 GIF 动图或截图)
案例背景:市民来电反映幸福家园小区南门路灯损坏,话务员完整记录了时间、地点(含参照物)、损坏数量及具体诉求。 AI 检测焦点:
- 完整性:自动比对录音中的“两盏”、“南门近超市”等细节,确认无遗漏。
- 一致性:确认话务员未歪曲市民的维修诉求。 AI 研判结果:
- 得分:100 分(优秀)
- 处置:高置信度 (High Confidence) -> 自动采信,无需人工干预。
案例背景:市民反映建设路共享单车乱停放,并在录音中反复强调“盲道被堵”且“险些造成盲人受伤”。工单仅记录“影响通行”。 AI 检测焦点:
- 完整性:识别出“盲道被堵”(重点治理项)和“安全隐患”(险些受伤)在工单中缺席。
- 风险意识:指出话务员未标记安全隐患,导致优先级评估偏低。 AI 研判结果:
- 得分:80 分(合格)
- 处置:中置信度 -> 建议人工复核。
- 修正建议:AI 自动补充“堵塞盲道”及“存在安全隐患”描述,并将优先级提升为“Urgent”。
案例背景:市民因化工厂异味问题多次投诉无果,情绪极度激动,扬言“要去拉横幅”、“找媒体曝光”,且提及“孩子住院”。工单仅记录为普通“异味反映”。 AI 检测焦点:
- 风险敏感性:捕捉到“拉横幅”(群体事件风险)、“找媒体”(舆情风险)及“孩子住院”(健康风险)。
- 一致性:判定话务员将“最后通牒”降级为“一般诉求”,属于严重失职。 AI 研判结果:
- 得分:45 分(存在风险)
- 处置:低置信度/高风险 -> 强制人工复核。
- 警示:系统标记为“严重漏报高危风险”,建议立即升级为“特急”工单。
案例背景:市民明确高喊“我要投诉烧烤店扰民”,话务员却在工单中将其包装为“市民咨询餐饮业经营政策”,试图通过“咨询件”规避“投诉件”的考核。 AI 检测焦点:
- 一致性:发现录音中的核心意图(投诉/维权)与工单定性(咨询/求助)存在根本性冲突。
- 处理判定:话务员试图“直接办结”该投诉,但 AI 判定此为线下扰民问题,必须“转办”至市场监管局。
- 性质判定:识别此类行为为恶劣的“指鹿为马”性质。 AI 研判结果:
- 得分:35 分(不合格)
- 处置:高置信度 -> 建议直接退回重写。
- 修正建议:处理方式修正为 “转办 (Dispatch)”,优先级修正为 “紧急”。
- 追责建议:系统明确指出该工单属于性质恶劣的定性篡改,建议追究话务员责任。
案例背景:市民举报公园内有人私搭乱建(违建),属于需要执法部门现场查处的投诉。话务员为了快速结案,将其作为“建议”记录并选择了“直接办结”。 AI 检测焦点:
- 流程合规:AI 识别出“违建拆除”属于行政执法范畴,严禁话务员自行办结。
- 空转风险:指出“直接办结”会导致工单无法流转至城管部门,形成无效工单。 AI 研判结果:
- 得分:75 分(合格,但存在流程硬伤)
- 处置:高置信度 -> 强制人工复核。
- 修正建议:处理方式修正为 “转办 (Dispatch)”,并建议转至城管局或园林局。
graph TD
User["用户 / 质检员"] -->|交互| Web["前端 (React + Vite)"]
Web -->|"HTTP POST"| Server["后端 (Express)"]
Server -->|"组装 Prompt"| LLM["Qwen3.6-Flash (大模型)"]
LLM -->|"返回 JSON"| Server
Server -->|"解析结果"| Web
Web -->|"可视化报告"| User
- 前端: React 19, TypeScript, Tailwind CSS 4, Lucide Icons, Vite 7
- 后端: Node.js, Express, OpenAI SDK (Adapter)
- AI 模型: Qwen3.6-Flash (via Aliyun DashScope)
- 提示词工程: 5层分层推理逻辑 (评分 -> 置信度 -> 策略 -> 校准 -> 修正)
我们提供了一键启动脚本,可自动安装依赖并启动服务:
./setup_and_run.sh首次运行前,请确保您已拥有 Node.js 环境。脚本会自动创建配置文件,请随后在 server/.env 中填入您的 API Key。
- Node.js (v18+)
- npm 或 yarn
- 阿里云 Qwen API Key (或兼容 OpenAI 格式的其他 LLM Key)
cd server
# 复制环境变量示例文件
cp .env.example .env
# 编辑 .env 文件,填入您的 QWEN_API_KEY
vim .env
npm install
node index.js后端默认运行在 http://localhost:3000
cd web
npm install
npm run dev前端默认运行在 http://localhost:5173
本项目支持通过 Cloudflare Pages 进行全栈部署,前端(Vite)和后端(Hono Functions)将运行在同一个域名下,无需跨域配置。
-
准备环境: 确保你已经安装了 wrangler CLI:
npm install -g wrangler
-
设置环境变量: 登录 Cloudflare Dashboard,进入你的 Pages 项目设置 -> Environment variables,添加以下变量:
QWEN_API_KEY: 你的阿里云 API KeyQWEN_BASE_URL:https://dashscope.aliyuncs.com/compatible-mode/v1QWEN_MODEL_NAME:qwen3.8-flashQWEN_ASR_MODEL:qwen-audio-3.0-asr-flash-streamQWEN_REALTIME_MODEL:paraformer-realtime-8k-v2
-
本地预览 (推荐): 在
web目录下运行以下命令,即可同时启动前端和后端:cd web npm install # 这一步会构建前端并启动 wrangler 本地环境 npm run build npx wrangler pages dev dist --binding QWEN_API_KEY=your_key
-
一键部署: 你可以直接通过命令行部署,或者连接 GitHub 仓库自动部署。
命令行部署:
cd web npm run build npx wrangler pages deploy dist --project-name govinsight-aiGitHub 自动部署 (推荐):
- 在 Cloudflare Pages 面板连接你的 GitHub 仓库。
- Build command:
npm run build - Build output directory:
dist - Root directory:
web(重要!因为前端代码在 web 目录下)
在 server/.env 中填写 API Key 后,需要重启服务才能生效:
-
使用一键脚本:
- 在终端按
Ctrl+C停止当前进程。 - 重新运行
./setup_and_run.sh。
- 在终端按
-
GitHub Codespaces:
- 同样建议停止并重启脚本。
- 或者直接使用
pkill node终止后台进程,然后重新运行启动命令。
- V0.1: 基础评分功能 (Basic Scoring)
- V0.2: 置信度评估与分级处置 (Confidence & Bucketing)
- V0.3: UI 重构、条件式修正生成、Mock 演示模式
- V0.3.2 (Latest): 一键自动化部署脚本、模型配置化、Node.js 运行时自动管理。
- V0.3.1: Dashboard 布局重构、评分标准 Tooltip、新 Logo 设计。
- V0.4: 检查工单直接办结与转办逻辑(Direct vs Dispatch)
- V0.5: 批量质检的调用和返回接口
- V0.6: 增加语音实时录入、音频上传转文字功能
- V0.7: 生成工单质检报告
- V1.0: 完整的仪表盘 (Dashboard) 与多租户支持
本项目采用 GNU GPL v3.0 许可证。
GovInsight-AI is an open-source intelligent quality inspection system powered by Large Language Models (LLM) (specifically Qwen3.6-Flash). It addresses the critical challenge of verification between "Call Transcripts" and "Operator Work Orders" in government service hotlines (e.g., 12345).
Traditional manual inspection is inefficient, inconsistent, and often fails to detect subtle semantic tampering. GovInsight-AI solves this by automatically comparing audio transcripts with work order records, accurately identifying missing key information, semantic deviations, and risk downgrading, while providing intelligent revision suggestions.
In the daily operation of government service hotlines (like 12345), the quality of work order records directly affects the efficiency of handling public appeals and citizen satisfaction. However, traditional manual quality inspection faces significant challenges:
- ⚡️ Low Efficiency: With massive volumes of calls and records, manual sampling rates are typically below 5%, leaving many problematic orders undetected.
- 📏 Inconsistent Standards: Subjective judgments vary greatly among different inspectors, making it difficult to form a unified and fair evaluation system.
- 🙈 Hidden Tampering: To avoid penalties, operators might privately change "Complaints" to "Consultations" or intentionally omit aggressive language, which is hard to verify without listening to every recording.
- 📉 Lagging Feedback: Inspections are usually post-event (T+1 or even T+7), making it impossible to intercept and correct errors before the work order is dispatched.
GovInsight-AI was born to solve these pain points by introducing LLM's semantic understanding capabilities into the inspection process, achieving full-volume, real-time, and objective intelligent detection.
GovInsight-AI is not just a scoring tool, but an Intelligent Assistant.
The system performs a deep scan of work orders based on four core dimensions:
- Completeness: Detects omission of key elements like time, location, involved parties, and specific demands.
- Consistency (Core Capability): Compares audio with the work order to find semantic tampering, factual deviations, or qualitative changes (e.g., turning a "Complaint" into a "Consultation").
- Clarity: Evaluates if the expression is clear, professional, and free of grammatical errors, ambiguity, or colloquialisms.
- Risk Awareness: Identifies if the operator ignored intense emotions, repeated complaint history, or potential risks of public opinion escalation.
Reject "Black Box" judgments! The system displays the AI's full reasoning process:
"The user explicitly mentioned 'this is the third complaint' in the recording, but this information was not recorded in the work order. This constitutes a key information omission and reduces the urgency of the issue..." This explainability allows inspectors to quickly verify and trust the AI's judgment.
Introducing a Confidence mechanism to categorize work orders into three types:
- ✅ Auto-Pass: Orders with Confidence ≥ 0.85 and no risks are automatically passed without human intervention.
- 👀 Sampling Review: Orders with Confidence between 0.70 - 0.84 enter the sampling pool.
- 🚨 Mandatory Review: Orders with Confidence < 0.70 or high risks (e.g., aggressive emotions) require mandatory human review.
When quality issues are detected, the AI not only reports errors but also automatically rewrites a standard work order. The system provides an intuitive Diff View, highlighting the differences between the original and the AI-suggested version, allowing operators or inspectors to adopt suggestions with one click.
(GIF or screenshots recommended here)
Context: A citizen reports a broken street light. The operator records the time, location, and issue accurately. AI Detection Focus:
- Completeness: Verifies details like "two lights" and "south gate near supermarket".
- Consistency: Confirms no distortion of the repair request. AI Verdict:
- Score: 100 (Excellent)
- Action: High Confidence -> Auto-Pass.
Context: A citizen reports shared bikes blocking the sidewalk, repeatedly emphasizing "blocking the blind lane" and "nearly causing injury to a blind person". The work order only records "bikes affecting traffic". AI Detection Focus:
- Completeness: Identifies missing critical details: "blocking blind lane" (priority issue) and "safety hazard".
- Risk Awareness: Flags the failure to mark the safety hazard. AI Verdict:
- Score: 80 (Qualified)
- Action: Medium Confidence -> Human Review Suggested.
- Revision: AI automatically adds "blocking blind lane" and "safety hazard", upgrading priority to "Urgent".
Context: A citizen complains about chemical odors for the 3rd time, threatening to "protest with banners" and mentioning "child hospitalized". The operator records it as a standard "odor complaint". AI Detection Focus:
- Risk Awareness: Captures high-risk keywords: "protest" (mass incident risk), "media exposure" (public opinion risk), and "child hospitalized" (health risk).
- Consistency: Determines the operator downgraded a "final ultimatum" to a "general request", a serious dereliction of duty. AI Verdict:
- Score: 45 (Risk)
- Action: Low Confidence / High Risk -> Mandatory Human Review.
- Alert: System flags "Serious Omission of High Risk", suggesting an immediate upgrade to "Emergency".
Context: A citizen explicitly shouts "I want to file a complaint about noise", but the operator records it as "Citizen consulting on catering policies" to avoid a complaint record. AI Detection Focus:
- Consistency: Detects a fundamental conflict between the core intent (Complaint) and work order type (Consultation).
- Nature Judgment: Identifies this as malicious "calling a stag a horse" (fact distortion). AI Verdict:
- Score: 35 (Unqualified)
- Action: High Confidence -> Reject & Rewrite.
- Accountability: System explicitly identifies malicious tampering and suggests accountability measures.
Context: A citizen reports illegal construction in a park, which requires on-site enforcement. To close the case quickly, the operator records it as a "Suggestion" and selects "Direct" (Direct Closure). AI Detection Focus:
- Process Compliance: AI identifies that "demolition of illegal construction" falls under administrative enforcement and strictly prohibits direct closure by the operator.
- Loop Risk: Points out that "Direct Closure" prevents the order from being dispatched to the Urban Management Department, resulting in an invalid order. AI Verdict:
- Score: 75 (Qualified, but with process flaws)
- Action: High Confidence -> Mandatory Human Review.
- Revision: Handling Type corrected to "Dispatch", suggesting transfer to the Urban Management Bureau or Parks Bureau.
graph TD
User["User / Inspector"] -->|Interaction| Web["Frontend (React + Vite)"]
Web -->|"HTTP POST"| Server["Backend (Express)"]
Server -->|"Construct Prompt"| LLM["Qwen3.5-Plus (LLM)"]
LLM -->|"Return JSON"| Server
Server -->|"Parse Result"| Web
Web -->|"Visual Report"| User
- Frontend: Built with React & Vite, providing an interactive dashboard for inspectors to view transcripts, work orders, and AI analysis results side-by-side.
- Backend: A lightweight Express server that handles API requests, constructs context-aware prompts (injecting history factors), and communicates with the LLM provider.
- Core Engine: Powered by Qwen3.5-Plus (via Aliyun DashScope), performing the 5-layer reasoning process to generate scores, confidence levels, and revisions.
- Frontend: React 19, TypeScript, Tailwind CSS 4, Lucide Icons, Vite 7
- Backend: Node.js, Express, OpenAI SDK (Adapter)
- AI Model: Qwen3.5-Plus (via Aliyun DashScope)
- Prompt Engineering: 5-layer reasoning logic (Scoring -> Confidence -> Strategy -> Calibration -> Revision)
- Backend:
cd server->cp .env.example .env->npm install->node index.js - Frontend:
cd web->npm install->npm run dev
GNU GPL v3.0 License