Dataset Viewer
Auto-converted to Parquet Duplicate
model
string
evaluation_date
string
task_id
string
agent_type
string
difficulty
string
prompt
string
success
bool
tool_called
bool
correct_tool
bool
final_answer_called
bool
response_correct
bool
tools_used
list
steps
int64
response
string
error
null
trace_id
string
execution_time_ms
float64
total_tokens
int64
cost_usd
float64
enhanced_trace_info
string
openai/gpt-5.4-nano
2026-04-24T11:23:23.860986
tool_weather_single
tool
easy
What's the weather in Paris, France?
true
true
true
true
true
[ "get_weather" ]
3
Paris, France: 20°C, Partly Cloudy.
null
0x7cd2b60fd98cacfa2ec4ade1b252e1e0
27,837.8496
9,858
0.001415
{"trace_id": "0x7cd2b60fd98cacfa2ec4ade1b252e1e0", "total_tokens": 9858, "duration_ms": 27837.849600000005, "cost_usd": 0.0014152, "span_count": 12}
openai/gpt-5.4-nano
2026-04-24T11:23:23.861029
tool_weather_compare
tool
medium
Compare the weather in Paris, France and London, UK. Which one is warmer?
true
true
true
true
true
[ "get_weather", "get_weather" ]
3
Paris is warmer right now: Paris is about 20°C (partly cloudy) versus London at about 15°C (rainy).
null
0xe0377777a7e65ba6060473a6a182d9ae
11,081.845672
10,215
0.001562
{"trace_id": "0xe0377777a7e65ba6060473a6a182d9ae", "total_tokens": 10215, "duration_ms": 11081.845671999998, "cost_usd": 0.0015615000000000002, "span_count": 13}
SMOLTRACE Logo

Tiny Agents. Total Visibility.

GitHub PyPI Documentation


SMOLTRACE Evaluation Results

This dataset contains evaluation results from a SMOLTRACE benchmark run.

Dataset Information

Field Value
Model openai/gpt-5.4-nano
Run ID lt-gpt54nano-e86e0118c1de4bf0
Agent Type tool
Total Tests 2
Generated 2026-04-24 11:23:26 UTC
Source Dataset kshitijthakkar/smoltrace-tasks

Schema

Column Type Description
model string Model identifier
evaluation_date string ISO timestamp of evaluation
task_id string Unique test case identifier
agent_type string "tool" or "code" agent type
difficulty string Test difficulty level
prompt string Test prompt/question
success bool Whether the test passed
tool_called bool Whether a tool was invoked
correct_tool bool Whether the correct tool was used
final_answer_called bool Whether final_answer was called
response_correct bool Whether the response was correct
tools_used string Comma-separated list of tools used
steps int Number of agent steps taken
response string Agent's final response
error string Error message if failed
trace_id string OpenTelemetry trace ID
execution_time_ms float Execution time in milliseconds
total_tokens int Total tokens consumed
cost_usd float API cost in USD
enhanced_trace_info string JSON with detailed trace data

Usage

from datasets import load_dataset

# Load the results dataset
ds = load_dataset("YOUR_USERNAME/smoltrace-results-TIMESTAMP")

# Filter successful tests
successful = ds.filter(lambda x: x['success'])

# Calculate success rate
success_rate = sum(1 for r in ds['train'] if r['success']) / len(ds['train']) * 100
print(f"Success Rate: {success_rate:.2f}%")

Related Datasets

This evaluation run also generated:

  • Traces Dataset: Detailed OpenTelemetry execution traces
  • Metrics Dataset: GPU utilization and environmental metrics
  • Leaderboard: Aggregated metrics for model comparison

About SMOLTRACE

SMOLTRACE is a comprehensive benchmarking and evaluation framework for Smolagents - HuggingFace's lightweight agent library.

Key Features

  • Automated agent evaluation with customizable test cases
  • OpenTelemetry-based tracing for detailed execution insights
  • GPU metrics collection (utilization, memory, temperature, power)
  • CO2 emissions and power cost tracking
  • Leaderboard aggregation and comparison

Quick Links

Installation

pip install smoltrace

Citation

If you use SMOLTRACE in your research, please cite:

@software{smoltrace,
  title = {SMOLTRACE: Benchmarking Framework for Smolagents},
  author = {Thakkar, Kshitij},
  url = {https://github.com/Mandark-droid/SMOLTRACE},
  year = {2025}
}

Generated by SMOLTRACE
Downloads last month
10