Skip to content

Starts a batch evaluation job that evaluates agent performance across multiple sessions

Description

Starts a batch evaluation job that evaluates agent performance across multiple sessions. Batch evaluations pull agent traces from CloudWatch Logs or an existing online evaluation configuration and run specified evaluators and insights against them.

Usage

bedrockagentcore_start_batch_evaluation(batchEvaluationName, evaluators,
  insights, dataSourceConfig, clientToken, evaluationMetadata, tags,
  kmsKeyArn, description, outputConfig)

Arguments

  • batchEvaluationName

    [required] The name of the batch evaluation. Must be unique within your account.

  • evaluators

    The list of evaluators to apply during the batch evaluation. Can include both built-in evaluators and custom evaluators. Maximum of 10 evaluators.

  • insights

    The list of insight analyses to run against sessions during the batch evaluation. Maximum of 10 insights.

  • dataSourceConfig

    [required] The data source configuration that specifies where to pull agent session traces from for evaluation.

  • clientToken

    A unique, case-sensitive identifier to ensure that the API request completes no more than one time. If this token matches a previous request, the service ignores the request, but does not return an error.

  • evaluationMetadata

    Optional metadata for the evaluation, including session-specific ground truth data and test scenario identifiers.

  • tags

    A map of tag keys and values to associate with the batch evaluation.

  • kmsKeyArn

    The ARN of the KMS key used to encrypt evaluation data. If provided, customer data is encrypted at rest with the specified key.

  • description

    The description of the batch evaluation.

  • outputConfig

    Output destination configuration.

Value

A list with the following syntax:

list(
  batchEvaluationId = "string",
  batchEvaluationArn = "string",
  batchEvaluationName = "string",
  evaluators = list(
    list(
      evaluatorId = "string"
    )
  ),
  insights = list(
    list(
      insightId = "string"
    )
  ),
  status = "PENDING"|"IN_PROGRESS"|"COMPLETED"|"COMPLETED_WITH_ERRORS"|"FAILED"|"STOPPING"|"STOPPED"|"DELETING",
  createdAt = as.POSIXct(
    "2015-01-01"
  ),
  outputConfig = list(
    cloudWatchConfig = list(
      logGroupName = "string",
      logStreamName = "string",
      metricsNamespace = "string",
      resultDestination = "DEDICATED_LOG_GROUP"|"SOURCE_LOG_GROUP"
    )
  ),
  tags = list(
    "string"
  ),
  kmsKeyArn = "string",
  description = "string"
)

Request syntax

svc$start_batch_evaluation(
  batchEvaluationName = "string",
  evaluators = list(
    list(
      evaluatorId = "string"
    )
  ),
  insights = list(
    list(
      insightId = "string"
    )
  ),
  dataSourceConfig = list(
    cloudWatchLogs = list(
      serviceNames = list(
        "string"
      ),
      logGroupNames = list(
        "string"
      ),
      logGroupNamePrefixes = list(
        "string"
      ),
      filterConfig = list(
        sessionIds = list(
          "string"
        ),
        timeRange = list(
          startTime = as.POSIXct(
            "2015-01-01"
          ),
          endTime = as.POSIXct(
            "2015-01-01"
          )
        ),
        sessionTraceIds = list(
          list(
            sessionId = "string",
            traceIds = list(
              "string"
            )
          )
        )
      )
    ),
    onlineEvaluationConfigSource = list(
      onlineEvaluationConfigArn = "string",
      timeRange = list(
        startTime = as.POSIXct(
          "2015-01-01"
        ),
        endTime = as.POSIXct(
          "2015-01-01"
        )
      )
    )
  ),
  clientToken = "string",
  evaluationMetadata = list(
    sessionMetadata = list(
      list(
        sessionId = "string",
        testScenarioId = "string",
        groundTruth = list(
          inline = list(
            assertions = list(
              list(
                text = "string"
              )
            ),
            expectedTrajectory = list(
              toolNames = list(
                "string"
              )
            ),
            turns = list(
              list(
                input = list(
                  prompt = "string"
                ),
                expectedResponse = list(
                  text = "string"
                )
              )
            )
          )
        ),
        metadata = list(
          "string"
        )
      )
    )
  ),
  tags = list(
    "string"
  ),
  kmsKeyArn = "string",
  description = "string",
  outputConfig = list(
    cloudWatchConfig = list(
      logGroupName = "string",
      logStreamName = "string",
      metricsNamespace = "string",
      resultDestination = "DEDICATED_LOG_GROUP"|"SOURCE_LOG_GROUP"
    )
  )
)