Starts a batch evaluation job that evaluates agent performance across multiple sessions¶
Description¶
Starts a batch evaluation job that evaluates agent performance across multiple sessions. Batch evaluations pull agent traces from CloudWatch Logs or an existing online evaluation configuration and run specified evaluators and insights against them.
Usage¶
bedrockagentcore_start_batch_evaluation(batchEvaluationName, evaluators,
insights, dataSourceConfig, clientToken, evaluationMetadata, tags,
kmsKeyArn, description, outputConfig)
Arguments¶
-
batchEvaluationName[required] The name of the batch evaluation. Must be unique within your account.
-
evaluatorsThe list of evaluators to apply during the batch evaluation. Can include both built-in evaluators and custom evaluators. Maximum of 10 evaluators.
-
insightsThe list of insight analyses to run against sessions during the batch evaluation. Maximum of 10 insights.
-
dataSourceConfig[required] The data source configuration that specifies where to pull agent session traces from for evaluation.
-
clientTokenA unique, case-sensitive identifier to ensure that the API request completes no more than one time. If this token matches a previous request, the service ignores the request, but does not return an error.
-
evaluationMetadataOptional metadata for the evaluation, including session-specific ground truth data and test scenario identifiers.
-
tagsA map of tag keys and values to associate with the batch evaluation.
-
kmsKeyArnThe ARN of the KMS key used to encrypt evaluation data. If provided, customer data is encrypted at rest with the specified key.
-
descriptionThe description of the batch evaluation.
-
outputConfigOutput destination configuration.
Value¶
A list with the following syntax:
list(
batchEvaluationId = "string",
batchEvaluationArn = "string",
batchEvaluationName = "string",
evaluators = list(
list(
evaluatorId = "string"
)
),
insights = list(
list(
insightId = "string"
)
),
status = "PENDING"|"IN_PROGRESS"|"COMPLETED"|"COMPLETED_WITH_ERRORS"|"FAILED"|"STOPPING"|"STOPPED"|"DELETING",
createdAt = as.POSIXct(
"2015-01-01"
),
outputConfig = list(
cloudWatchConfig = list(
logGroupName = "string",
logStreamName = "string",
metricsNamespace = "string",
resultDestination = "DEDICATED_LOG_GROUP"|"SOURCE_LOG_GROUP"
)
),
tags = list(
"string"
),
kmsKeyArn = "string",
description = "string"
)
Request syntax¶
svc$start_batch_evaluation(
batchEvaluationName = "string",
evaluators = list(
list(
evaluatorId = "string"
)
),
insights = list(
list(
insightId = "string"
)
),
dataSourceConfig = list(
cloudWatchLogs = list(
serviceNames = list(
"string"
),
logGroupNames = list(
"string"
),
logGroupNamePrefixes = list(
"string"
),
filterConfig = list(
sessionIds = list(
"string"
),
timeRange = list(
startTime = as.POSIXct(
"2015-01-01"
),
endTime = as.POSIXct(
"2015-01-01"
)
),
sessionTraceIds = list(
list(
sessionId = "string",
traceIds = list(
"string"
)
)
)
)
),
onlineEvaluationConfigSource = list(
onlineEvaluationConfigArn = "string",
timeRange = list(
startTime = as.POSIXct(
"2015-01-01"
),
endTime = as.POSIXct(
"2015-01-01"
)
)
)
),
clientToken = "string",
evaluationMetadata = list(
sessionMetadata = list(
list(
sessionId = "string",
testScenarioId = "string",
groundTruth = list(
inline = list(
assertions = list(
list(
text = "string"
)
),
expectedTrajectory = list(
toolNames = list(
"string"
)
),
turns = list(
list(
input = list(
prompt = "string"
),
expectedResponse = list(
text = "string"
)
)
)
)
),
metadata = list(
"string"
)
)
)
),
tags = list(
"string"
),
kmsKeyArn = "string",
description = "string",
outputConfig = list(
cloudWatchConfig = list(
logGroupName = "string",
logStreamName = "string",
metricsNamespace = "string",
resultDestination = "DEDICATED_LOG_GROUP"|"SOURCE_LOG_GROUP"
)
)
)