Create Model Evaluation Task
Access Entry
On the model details page, click the Model Evaluation button to navigate to the evaluation task creation page.
Note
Only some models support creating evaluation tasks. If the desired model does not have a "Model Evaluation" option, contact the platform administrator.
Configuration Parameters
On the model evaluation task creation page, fill in the following configuration, then click Create Evaluation:
| Parameter | Description |
|---|---|
| Task Name | Custom evaluation task name |
| Models | Platform model IDs; compare up to 3 models in one task |
| Description | Optional note for this evaluation |
| Datasets | System recommended: benchmarks available in the selected framework image; Custom: existing repositories on the platform. See Custom Evaluation Datasets for formats |
| Region / Resource | Cluster and compute specification |
| Evaluation Framework | Choose the framework (OpenCompass, EvalScope, or lm-evaluation-harness), then the framework version |
| Framework args / vLLM args | Shown when the selected version declares engine_args. These change runtime behavior, not the scoring formula |
| Advanced options | Collapsed. Prompt template and scoring plugins default from the selected dataset and framework. Overrides are not leaderboard-comparable |
For parameter boundaries, when to change them, and the report snapshot, see Evaluation Custom Parameters. For Prompt and plugin details, see Scoring Plugins and Prompt Templates and Metrics Configuration.
View Evaluation Results
After creation, use the top navigation to open Model Training & Evaluation → Model Evaluation to view the running status and results of all evaluation tasks. You can also view them centrally in Resource Management.