# Main classes

## EvaluationModuleInfo[[evaluate.EvaluationModuleInfo]]

The base class `EvaluationModuleInfo` implements a the logic for the subclasses `MetricInfo`, `ComparisonInfo`, and `MeasurementInfo`.

#### evaluate.EvaluationModuleInfo[[evaluate.EvaluationModuleInfo]]

```python
evaluate.EvaluationModuleInfo(description: str, citation: str, features: typing.Union[datasets.features.features.Features, typing.List[datasets.features.features.Features]], inputs_description: str = <factory>, homepage: str = <factory>, license: str = <factory>, codebase_urls: typing.List[str] = <factory>, reference_urls: typing.List[str] = <factory>, streamable: bool = False, format: typing.Optional[str] = None, module_type: str = 'metric', module_name: typing.Optional[str] = None, config_name: typing.Optional[str] = None, experiment_id: typing.Optional[str] = None)
```

[Source](https://github.com/huggingface/evaluate/blob/main/src/evaluate/info.py#L35)

Base class to store information about an evaluation used for `MetricInfo`, `ComparisonInfo`,
and `MeasurementInfo`.

`EvaluationModuleInfo` documents an evaluation, including its name, version, and features.
See the constructor arguments and properties for a full list.

Note: Not all fields are known on construction and may be updated later.

#### from_directory[[evaluate.EvaluationModuleInfo.from_directory]]

```python
from_directory(metric_info_dir)
```

[Source](https://github.com/huggingface/evaluate/blob/main/src/evaluate/info.py#L92)

**Parameters:**

metric_info_dir (`str`) : The directory containing the `metric_info` JSON file. This should be the root directory of a specific metric version.

Create `EvaluationModuleInfo` from the JSON file in `metric_info_dir`.

Example:

```py
>>> my_metric = EvaluationModuleInfo.from_directory("/path/to/directory/")
```

#### write_to_directory[[evaluate.EvaluationModuleInfo.write_to_directory]]

```python
write_to_directory(metric_info_dir)
```

[Source](https://github.com/huggingface/evaluate/blob/main/src/evaluate/info.py#L72)

**Parameters:**

metric_info_dir (`str`) : The directory to save `metric_info_dir` to.

Write `EvaluationModuleInfo` as JSON to `metric_info_dir`.
Also save the license separately in LICENSE.

Example:

```py
>>> my_metric.info.write_to_directory("/path/to/directory/")
```

#### evaluate.MetricInfo[[evaluate.MetricInfo]]

```python
evaluate.MetricInfo(description: str, citation: str, features: typing.Union[datasets.features.features.Features, typing.List[datasets.features.features.Features]], inputs_description: str = <factory>, homepage: str = <factory>, license: str = <factory>, codebase_urls: typing.List[str] = <factory>, reference_urls: typing.List[str] = <factory>, streamable: bool = False, format: typing.Optional[str] = None, module_type: str = 'metric', module_name: typing.Optional[str] = None, config_name: typing.Optional[str] = None, experiment_id: typing.Optional[str] = None)
```

[Source](https://github.com/huggingface/evaluate/blob/main/src/evaluate/info.py#L122)

Information about a metric.

`EvaluationModuleInfo` documents a metric, including its name, version, and features.
See the constructor arguments and properties for a full list.

Note: Not all fields are known on construction and may be updated later.

#### evaluate.ComparisonInfo[[evaluate.ComparisonInfo]]

```python
evaluate.ComparisonInfo(description: str, citation: str, features: typing.Union[datasets.features.features.Features, typing.List[datasets.features.features.Features]], inputs_description: str = <factory>, homepage: str = <factory>, license: str = <factory>, codebase_urls: typing.List[str] = <factory>, reference_urls: typing.List[str] = <factory>, streamable: bool = False, format: typing.Optional[str] = None, module_type: str = 'comparison', module_name: typing.Optional[str] = None, config_name: typing.Optional[str] = None, experiment_id: typing.Optional[str] = None)
```

[Source](https://github.com/huggingface/evaluate/blob/main/src/evaluate/info.py#L135)

Information about a comparison.

`EvaluationModuleInfo` documents a comparison, including its name, version, and features.
See the constructor arguments and properties for a full list.

Note: Not all fields are known on construction and may be updated later.

#### evaluate.MeasurementInfo[[evaluate.MeasurementInfo]]

```python
evaluate.MeasurementInfo(description: str, citation: str, features: typing.Union[datasets.features.features.Features, typing.List[datasets.features.features.Features]], inputs_description: str = <factory>, homepage: str = <factory>, license: str = <factory>, codebase_urls: typing.List[str] = <factory>, reference_urls: typing.List[str] = <factory>, streamable: bool = False, format: typing.Optional[str] = None, module_type: str = 'measurement', module_name: typing.Optional[str] = None, config_name: typing.Optional[str] = None, experiment_id: typing.Optional[str] = None)
```

[Source](https://github.com/huggingface/evaluate/blob/main/src/evaluate/info.py#L148)

Information about a measurement.

`EvaluationModuleInfo` documents a measurement, including its name, version, and features.
See the constructor arguments and properties for a full list.

Note: Not all fields are known on construction and may be updated later.

## EvaluationModule[[evaluate.EvaluationModule]]

The base class `EvaluationModule` implements a the logic for the subclasses `Metric`, `Comparison`, and `Measurement`.

#### evaluate.EvaluationModule[[evaluate.EvaluationModule]]

```python
evaluate.EvaluationModule(config_name: typing.Optional[str] = None, keep_in_memory: bool = False, cache_dir: typing.Optional[str] = None, num_process: int = 1, process_id: int = 0, seed: typing.Optional[int] = None, experiment_id: typing.Optional[str] = None, hash: str = None, max_concurrent_cache_files: int = 10000, timeout: typing.Union[int, float] = 100, **kwargs)
```

[Source](https://github.com/huggingface/evaluate/blob/main/src/evaluate/module.py#L149)

**Parameters:**

config_name (`str`) : This is used to define a hash specific to a module computation script and prevents the module's data to be overridden when the module loading script is modified.

keep_in_memory (`bool`) : Keep all predictions and references in memory. Not possible in distributed settings.

cache_dir (`str`) : Path to a directory in which temporary prediction/references data will be stored. The data directory should be located on a shared file-system in distributed setups.

num_process (`int`) : Specify the total number of nodes in a distributed settings. This is useful to compute module in distributed setups (in particular non-additive modules like F1).

process_id (`int`) : Specify the id of the current process in a distributed setup (between 0 and num_process-1) This is useful to compute module in distributed setups (in particular non-additive metrics like F1).

seed (`int`, optional) : If specified, this will temporarily set numpy's random seed when [compute()](/docs/evaluate/main/en/package_reference/main_classes#evaluate.EvaluationModule.compute) is run.

experiment_id (`str`) : A specific experiment id. This is used if several distributed evaluations share the same file system. This is useful to compute module in distributed setups (in particular non-additive metrics like F1).

hash (`str`) : Used to identify the evaluation module according to the hashed file contents.

max_concurrent_cache_files (`int`) : Max number of concurrent module cache files (default `10000`).

timeout (`Union[int, float]`) : Timeout in second for distributed setting synchronization.

A `EvaluationModule` is the base class and common API for metrics, comparisons, and measurements.

#### add[[evaluate.EvaluationModule.add]]

```python
add(prediction = None, reference = None, **kwargs)
```

[Source](https://github.com/huggingface/evaluate/blob/main/src/evaluate/module.py#L548)

**Parameters:**

prediction (`list/array/tensor`, *optional*) : Predictions.

reference (`list/array/tensor`, *optional*) : References.

Add one prediction and reference for the evaluation module's stack.

Example:

```py
>>> import evaluate
>>> accuracy = evaluate.load("accuracy")
>>> accuracy.add(references=[0,1], predictions=[1,0])
```

#### add_batch[[evaluate.EvaluationModule.add_batch]]

```python
add_batch(predictions = None, references = None, **kwargs)
```

[Source](https://github.com/huggingface/evaluate/blob/main/src/evaluate/module.py#L488)

**Parameters:**

predictions (`list/array/tensor`, *optional*) : Predictions.

references (`list/array/tensor`, *optional*) : References.

Add a batch of predictions and references for the evaluation module's stack.

Example:

```py
>>> import evaluate
>>> accuracy = evaluate.load("accuracy")
>>> for refs, preds in zip([[0,1],[0,1]], [[1,0],[0,1]]):
...     accuracy.add_batch(references=refs, predictions=preds)
```

#### compute[[evaluate.EvaluationModule.compute]]

```python
compute(predictions = None, references = None, **kwargs)
```

[Source](https://github.com/huggingface/evaluate/blob/main/src/evaluate/module.py#L415)

**Parameters:**

predictions (`list/array/tensor`, *optional*) : Predictions.

references (`list/array/tensor`, *optional*) : References.

- ****kwargs** (optional) : Keyword arguments that will be forwarded to the evaluation module [compute()](/docs/evaluate/main/en/package_reference/main_classes#evaluate.EvaluationModule.compute) method (see details in the docstring).

**Returns:** `dict` or `None`

- Dictionary with the results if this evaluation module is run on the main process (`process_id == 0`).
- `None` if the evaluation module is not run on the main process (`process_id != 0`).

Compute the evaluation module.

Usage of positional arguments is not allowed to prevent mistakes.

```py
>>> import evaluate
>>> accuracy =  evaluate.load("accuracy")
>>> accuracy.compute(predictions=[0, 1, 1, 0], references=[0, 1, 0, 1])
```

#### download_and_prepare[[evaluate.EvaluationModule.download_and_prepare]]

```python
download_and_prepare(download_config: typing.Optional[datasets.download.download_config.DownloadConfig] = None, dl_manager: typing.Optional[datasets.download.download_manager.DownloadManager] = None)
```

[Source](https://github.com/huggingface/evaluate/blob/main/src/evaluate/module.py#L677)

**Parameters:**

download_config (`DownloadConfig`, *optional*) : Specific download configuration parameters.

dl_manager (`DownloadManager`, *optional*) : Specific download manager to use.

Downloads and prepares evaluation module for reading.

Example:

```py
>>> import evaluate
```

#### evaluate.Metric[[evaluate.Metric]]

```python
evaluate.Metric(config_name: typing.Optional[str] = None, keep_in_memory: bool = False, cache_dir: typing.Optional[str] = None, num_process: int = 1, process_id: int = 0, seed: typing.Optional[int] = None, experiment_id: typing.Optional[str] = None, hash: str = None, max_concurrent_cache_files: int = 10000, timeout: typing.Union[int, float] = 100, **kwargs)
```

[Source](https://github.com/huggingface/evaluate/blob/main/src/evaluate/module.py#L782)

**Parameters:**

config_name (`str`) : This is used to define a hash specific to a metric computation script and prevents the metric's data to be overridden when the metric loading script is modified.

keep_in_memory (`bool`) : Keep all predictions and references in memory. Not possible in distributed settings.

cache_dir (`str`) : Path to a directory in which temporary prediction/references data will be stored. The data directory should be located on a shared file-system in distributed setups.

num_process (`int`) : Specify the total number of nodes in a distributed settings. This is useful to compute metrics in distributed setups (in particular non-additive metrics like F1).

process_id (`int`) : Specify the id of the current process in a distributed setup (between 0 and num_process-1) This is useful to compute metrics in distributed setups (in particular non-additive metrics like F1).

seed (`int`, *optional*) : If specified, this will temporarily set numpy's random seed when [compute()](/docs/evaluate/main/en/package_reference/main_classes#evaluate.EvaluationModule.compute) is run.

experiment_id (`str`) : A specific experiment id. This is used if several distributed evaluations share the same file system. This is useful to compute metrics in distributed setups (in particular non-additive metrics like F1).

max_concurrent_cache_files (`int`) : Max number of concurrent metric cache files (default `10000`).

timeout (`Union[int, float]`) : Timeout in second for distributed setting synchronization.

A Metric is the base class and common API for all metrics.

#### evaluate.Comparison[[evaluate.Comparison]]

```python
evaluate.Comparison(config_name: typing.Optional[str] = None, keep_in_memory: bool = False, cache_dir: typing.Optional[str] = None, num_process: int = 1, process_id: int = 0, seed: typing.Optional[int] = None, experiment_id: typing.Optional[str] = None, hash: str = None, max_concurrent_cache_files: int = 10000, timeout: typing.Union[int, float] = 100, **kwargs)
```

[Source](https://github.com/huggingface/evaluate/blob/main/src/evaluate/module.py#L812)

**Parameters:**

config_name (`str`) : This is used to define a hash specific to a comparison computation script and prevents the comparison's data to be overridden when the comparison loading script is modified.

keep_in_memory (`bool`) : Keep all predictions and references in memory. Not possible in distributed settings.

cache_dir (`str`) : Path to a directory in which temporary prediction/references data will be stored. The data directory should be located on a shared file-system in distributed setups.

num_process (`int`) : Specify the total number of nodes in a distributed settings. This is useful to compute  comparisons in distributed setups (in particular non-additive comparisons).

process_id (`int`) : Specify the id of the current process in a distributed setup (between 0 and num_process-1) This is useful to compute  comparisons in distributed setups (in particular non-additive comparisons).

seed (`int`, *optional*) : If specified, this will temporarily set numpy's random seed when [compute()](/docs/evaluate/main/en/package_reference/main_classes#evaluate.EvaluationModule.compute) is run.

experiment_id (`str`) : A specific experiment id. This is used if several distributed evaluations share the same file system. This is useful to compute  comparisons in distributed setups (in particular non-additive comparisons).

max_concurrent_cache_files (`int`) : Max number of concurrent comparison cache files (default `10000`).

timeout (`Union[int, float]`) : Timeout in second for distributed setting synchronization.

A Comparison is the base class and common API for all comparisons.

#### evaluate.Measurement[[evaluate.Measurement]]

```python
evaluate.Measurement(config_name: typing.Optional[str] = None, keep_in_memory: bool = False, cache_dir: typing.Optional[str] = None, num_process: int = 1, process_id: int = 0, seed: typing.Optional[int] = None, experiment_id: typing.Optional[str] = None, hash: str = None, max_concurrent_cache_files: int = 10000, timeout: typing.Union[int, float] = 100, **kwargs)
```

[Source](https://github.com/huggingface/evaluate/blob/main/src/evaluate/module.py#L842)

**Parameters:**

config_name (`str`) : This is used to define a hash specific to a measurement computation script and prevents the measurement's data to be overridden when the measurement loading script is modified.

keep_in_memory (`bool`) : Keep all predictions and references in memory. Not possible in distributed settings.

cache_dir (`str`) : Path to a directory in which temporary prediction/references data will be stored. The data directory should be located on a shared file-system in distributed setups.

num_process (`int`) : Specify the total number of nodes in a distributed settings. This is useful to compute measurements in distributed setups (in particular non-additive measurements).

process_id (`int`) : Specify the id of the current process in a distributed setup (between 0 and num_process-1) This is useful to compute measurements in distributed setups (in particular non-additive measurements).

seed (`int`, *optional*) : If specified, this will temporarily set numpy's random seed when [compute()](/docs/evaluate/main/en/package_reference/main_classes#evaluate.EvaluationModule.compute) is run.

experiment_id (`str`) : A specific experiment id. This is used if several distributed evaluations share the same file system. This is useful to compute measurements in distributed setups (in particular non-additive measurements).

max_concurrent_cache_files (`int`) : Max number of concurrent measurement cache files (default `10000`).

timeout (`Union[int, float]`) : Timeout in second for distributed setting synchronization.

A Measurement is the base class and common API for all measurements.

## CombinedEvaluations[[evaluate.combine]]

The `combine` function allows to combine multiple `EvaluationModule`s into a single `CombinedEvaluations`.

#### evaluate.combine[[evaluate.combine]]

```python
evaluate.combine(evaluations, force_prefix = False)
```

[Source](https://github.com/huggingface/evaluate/blob/main/src/evaluate/module.py#L1008)

**Parameters:**

evaluations (`Union[list, dict]`) : A list or dictionary of evaluation modules. The modules can either be passed as strings or loaded `EvaluationModule`s. If a dictionary is passed its keys are the names used and the values the modules. The names are used as prefix in case there are name overlaps in the returned results of each module or if `force_prefix=True`.

force_prefix (`bool`, *optional*, defaults to `False`) : If `True` all scores from the modules are prefixed with their name. If a dictionary is passed the keys are used as name otherwise the module's name.

Combines several metrics, comparisons, or measurements into a single `CombinedEvaluations` object that
can be used like a single evaluation module.

If two scores have the same name, then they are prefixed with their module names.
And if two modules have the same name, please use a dictionary to give them different names, otherwise an integer id is appended to the prefix.

Examples:

```py
>>> import evaluate
>>> accuracy = evaluate.load("accuracy")
>>> f1 = evaluate.load("f1")
>>> clf_metrics = combine(["accuracy", "f1"])
```

#### evaluate.CombinedEvaluations[[evaluate.CombinedEvaluations]]

```python
evaluate.CombinedEvaluations(evaluation_modules, force_prefix = False)
```

[Source](https://github.com/huggingface/evaluate/blob/main/src/evaluate/module.py#L872)

#### add[[evaluate.CombinedEvaluations.add]]

```python
add(prediction = None, reference = None, **kwargs)
```

[Source](https://github.com/huggingface/evaluate/blob/main/src/evaluate/module.py#L895)

**Parameters:**

predictions (`list/array/tensor`, *optional*) : Predictions.

references (`list/array/tensor`, *optional*) : References.

Add one prediction and reference for each evaluation module's stack.

Example:

```py
>>> import evaluate
>>> accuracy = evaluate.load("accuracy")
>>> f1 = evaluate.load("f1")
>>> clf_metrics = combine(["accuracy", "f1"])
>>> for ref, pred in zip([0,1,0,1], [1,0,0,1]):
...     clf_metrics.add(references=ref, predictions=pred)
```

#### add_batch[[evaluate.CombinedEvaluations.add_batch]]

```python
add_batch(predictions = None, references = None, **kwargs)
```

[Source](https://github.com/huggingface/evaluate/blob/main/src/evaluate/module.py#L920)

**Parameters:**

predictions (`list/array/tensor`, *optional*) : Predictions.

references (`list/array/tensor`, *optional*) : References.

Add a batch of predictions and references for each evaluation module's stack.

Example:
```py
>>> import evaluate
>>> accuracy = evaluate.load("accuracy")
>>> f1 = evaluate.load("f1")
>>> clf_metrics = combine(["accuracy", "f1"])
>>> for refs, preds in zip([[0,1],[0,1]], [[1,0],[0,1]]):
...     clf_metrics.add(references=refs, predictions=preds)
```

#### compute[[evaluate.CombinedEvaluations.compute]]

```python
compute(predictions = None, references = None, **kwargs)
```

[Source](https://github.com/huggingface/evaluate/blob/main/src/evaluate/module.py#L944)

**Parameters:**

predictions (`list/array/tensor`, *optional*) : Predictions.

references (`list/array/tensor`, *optional*) : References.

- ****kwargs** (*optional*) : Keyword arguments that will be forwarded to the evaluation module [compute()](/docs/evaluate/main/en/package_reference/main_classes#evaluate.EvaluationModule.compute) method (see details in the docstring).

**Returns:** `dict` or `None`

- Dictionary with the results if this evaluation module is run on the main process (`process_id == 0`).
- `None` if the evaluation module is not run on the main process (`process_id != 0`).

Compute each evaluation module.

Usage of positional arguments is not allowed to prevent mistakes.

Example:

```py
>>> import evaluate
>>> accuracy = evaluate.load("accuracy")
>>> f1 = evaluate.load("f1")
>>> clf_metrics = combine(["accuracy", "f1"])
>>> clf_metrics.compute(predictions=[0,1], references=[1,1])
{'accuracy': 0.5, 'f1': 0.6666666666666666}
```

