Pinned Loading
-
qwen3-financial-sft
qwen3-financial-sft PublicTRL + PEFT LoRA/QLoRA supervised fine-tuning of Qwen3-4B-Instruct for financial numeric QA (TAT-QA/FinQA-style) over hybrid table+text reports, with MLflow tracking and an offline evaluation harness.
Python
-
qwen3-financial-sft-unsloth-dpo
qwen3-financial-sft-unsloth-dpo PublicUnsloth QLoRA SFT then DPO on Qwen3-4B for financial QA — MLflow-tracked, type-aware evaluation surfacing a real SFT-vs-DPO tradeoff the aggregate score hid.
Python
-
policy-gradient-from-scratch
policy-gradient-from-scratch PublicPolicy-gradient RL built up from REINFORCE to GRPO in plain PyTorch, no RL library — makes the one shared idea across the algorithm family visible, and what each successive method actually adds.
Python
-
kafka-grid-intelligence
kafka-grid-intelligence PublicReal-time GB grid-stress prediction over Kafka, with a model-skill audit (persistence vs periodicity)
HTML
-
nhs-care-access-agent
nhs-care-access-agent PublicEvaluation-first, model-agnostic NHS care-access agent over the nhs-intelligence-mcp tool layer
Python
-
If the problem persists, check the GitHub status page or contact support.

