You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
* Revise example prompt and answer to English
Updated example prompt and answer in README with English content.
* Fix typo issue and add link to generative ai readme.md
Copy file name to clipboardExpand all lines: generative_ai/README.md
+2-1Lines changed: 2 additions & 1 deletion
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -6,7 +6,8 @@ These examples showcases Amazon SageMaker's capabilities in the exciting field o
6
6
7
7
-[Fine-tuning and deploying a Hugging Face summarization model on SageMaker with your own scripts and dataset](sm-finetuning_huggingface_with_your_own_scripts_and_data/sm-finetuning_huggingface_with_your_own_scripts_and_data.ipynb)
8
8
-[Fine-tuning and deploying the Mixtral 8x7B LLM In SageMaker with Hugging Face, using QLoRA Parameter-Efficient Fine-Tuning](sm-mixtral_8x7b_fine_tune_and_deploy/sm-mixtral_8x7b_fine_tune_and_deploy.ipynb)
9
-
-[Serve large models on SageMaker with DeepSpeed Container](sm-djl_deepspeed_bloom_176b_deploy.ipynb)
9
+
-[Qwen 8B LLM Fine-tuning with SFT and GRPO, and Deployment on AWS SageMaker](sm-qwen3_8b_fine_tune_and_deploy/README.md)
10
+
-[Serve large models on SageMaker with DeepSpeed Container](sm-mixtral_8x7b_fine_tune_and_deploy/sm-mixtral_8x7b_fine_tune_and_deploy.ipynb)(sm-djl_deepspeed_bloom_176b_deploy.ipynb)
10
11
-[Accelerate SageMaker-PyTorch FSDP Training of Llama-v2 (or GPT-NeoX) with FP8 on P5 instances](sm-fsdp_training_of_llama_v2_with_fp8_on_p5.ipynb)
11
12
-[Fine-tune Code Llama, Deploy and Evaluate the Fine-tuning with Human-eval Repository](sm-jumpstart_foundation_code_llama_fine_tuning_human_eval.ipynb)
12
13
-[SageMaker JumpStart Foundation Models - Fine-tuning text generation GPT-J 6B model on domain specific dataset](sm-jumpstart_foundation_finetuning_gpt_j_6b_domain_adaptation.ipynb)
|`formatting`| 5% | Correct output format (9 categories) |
634
634
635
635
To customize the reward function, edit `2_trainning_grpo/docker/reward_function/math.py`:
636
636
637
637
```python
638
638
def compute_score(
639
639
reward_inputs: list[dict[str, Any]],
640
-
recall_weight: float = 0.35,
641
-
precision_weight: float = 0.2,
642
-
accuracy_weight: float = 0.35,
643
-
match_quality_weight: float = 0.05,
640
+
recall_weight: float = 0.30,
641
+
precision_weight: float = 0.25,
642
+
accuracy_weight: float = 0.30,
643
+
match_quality_weight: float = 0.10,
644
644
formatting_weight: float = 0.05
645
645
) -> list[dict[str, float]]:
646
646
"""
@@ -683,6 +683,7 @@ GRPO training uses **Parquet format** with prompt-answer pairs. Upload your data
683
683
|`answer`| string | The expected ground truth response (used by reward function) |
684
684
685
685
**Example row:**
686
+
686
687
| Column | Content |
687
688
|--------|---------|
688
689
|`problem`| You are a professional and rigorous product tagging expert, responsible for automatically generating and classifying tags based on the provided product information...<br><br>Product Name: MI Amazon Usb Type-C Cable Smartphone Charging (Black) \|Connectivity: Usb 2.0 (Sync And Charging)\| Universal For All Type-C Devices (Grey)... |
0 commit comments