> ## Documentation Index
> Fetch the complete documentation index at: https://docs.bespokelabs.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Aspect-based sentiment analysis

> Distill GPT-4o into a small Llama model for restaurant review sentiment, using Curator and Together fine-tuning.

This example uses Curator to move a skill from a large model to a much smaller 8B parameter model. GPT-4o labels the sentiment of Yelp restaurant reviews, and those labels are used to fine-tune Llama 3.1 8B Instruct with the Together fine-tuning API.

You can also run this example in [Colab](https://colab.research.google.com/drive/1W2jf6v-ZwQ7mku1pcfdIMOawKY99Shi9).

The model reads a review and returns a sentiment for each aspect of the restaurant. Take this review as an example.

```text theme={null}
The food was good, but the service was slow.
```

It returns this JSON. The full output also has `ambience_sentiment`, `price_sentiment`, and `overall_sentiment`.

```json theme={null}
{
    "food_sentiment": "Positive",
    "service_sentiment": "Negative"
}
```

## Setup

```bash theme={null}
pip install bespokelabs-curator datasets together
```

```python theme={null}
import getpass
import json
import os

from datasets import load_dataset
from together import Together

from bespokelabs import curator

os.environ["OPENAI_API_KEY"] = getpass.getpass("Enter your OpenAI API key: ")
os.environ["TOGETHER_API_KEY"] = getpass.getpass("Enter your Together API key: ")
# The Curator Viewer lets you look at the data. Remove this line to turn it off.
os.environ["CURATOR_VIEWER"] = "1"
```

## Label the data

The prompt asks the model to rate five aspects of a review. This example does not use structured output, because the same `LLM` class also runs the small base model later, and many small models do not support structured output. The prompt asks for JSON in a code block instead.

````python theme={null}
PROMPT = """You are a sentiment analysis expert specializing in restaurant reviews. You need to analyze the sentiment of the given restaurant review.

Analyze the review for the following specific aspects:
1. Food: Quality, taste, presentation, menu variety, etc.
2. Service: Staff behavior, responsiveness, professionalism, etc.
3. Ambience: Atmosphere, decor, comfort, noise level, etc.
4. Price: Value for money, affordability, etc.
5. Overall: General impression of the restaurant experience

For each aspect, classify the sentiment as exactly one of the following:
- Positive: The review expresses satisfaction or praise
- Negative: The review expresses dissatisfaction or criticism
- Neutral: The review is balanced or doesn't mention the aspect

If an aspect is not mentioned in the review, classify it as Neutral.

Output the sentiment for each aspect in the following format:
```json
{
    "food_sentiment": "Positive",
    "service_sentiment": "Negative",
    "ambience_sentiment": "Neutral",
    "price_sentiment": "Positive",
    "overall_sentiment": "Negative"
}```
"""

ASPECTS = ["food_sentiment", "service_sentiment", "ambience_sentiment", "price_sentiment", "overall_sentiment"]


class AspectBasedSentimentCurator(curator.LLM):
    def prompt(self, input: dict) -> list:
        return [
            {"role": "system", "content": PROMPT},
            {"role": "user", "content": f"The review is: {input['text']}"},
        ]

    def parse(self, input: dict, raw_response: str) -> dict:
        try:
            response = json.loads(raw_response.split("```json")[1].split("```")[0])
        except (IndexError, json.JSONDecodeError):
            response = {}
        return {**input, **{aspect: response.get(aspect, "None") for aspect in ASPECTS}}
````

Load the Yelp reviews. You can upload them to the Curator Viewer to look through them first.

```python theme={null}
source_dataset = load_dataset("bespokelabs/yelp_restaurant_reviews", split="train")

url = curator.push_to_viewer(source_dataset)
```

Label the reviews with GPT-4o.

```python theme={null}
annotated_dataset = AspectBasedSentimentCurator(
    model_name="gpt-4o",
    generation_params={"temperature": 0.0},
)(source_dataset).dataset
```

## Split the data

Rename the label columns with a `_gt` suffix, for ground truth. Then use 90% of the rows for training and 10% for testing.

```python theme={null}
for aspect in ASPECTS:
    annotated_dataset = annotated_dataset.rename_column(aspect, f"{aspect}_gt")

split = int(len(annotated_dataset) * 0.9)
train_dataset = annotated_dataset.select(range(split))
test_dataset = annotated_dataset.select(range(split, len(annotated_dataset)))
```

## Evaluate the base model

This function compares the model output with the GPT-4o labels and returns the accuracy for each aspect.

```python theme={null}
def evaluate_sentiment(dataset):
    """Compare model output with ground truth for each aspect."""
    aspect_accuracies = {}
    for aspect in ASPECTS:
        correct = sum(1 for i in range(len(dataset)) if dataset[aspect][i] == dataset[f"{aspect}_gt"][i])
        aspect_accuracies[aspect] = correct / len(dataset) if len(dataset) > 0 else 0

    overall_accuracy = sum(aspect_accuracies.values()) / len(aspect_accuracies)
    return {"overall_accuracy": overall_accuracy, "aspect_accuracies": aspect_accuracies}
```

Run the base Llama 3.1 8B model on the test set through Together and LiteLLM.

```python theme={null}
small_model_output = AspectBasedSentimentCurator(
    model_name="together_ai/meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo",
    backend="litellm",
    generation_params={"temperature": 0.0},
    backend_params={"max_tokens_per_minute": 100_000_000},
)(test_dataset).dataset

base_eval = evaluate_sentiment(small_model_output)
print(json.dumps(base_eval, indent=4))
```

```json theme={null}
{
    "overall_accuracy": 0.8271653543307087,
    "aspect_accuracies": {
        "food_sentiment": 0.8661417322834646,
        "service_sentiment": 0.9251968503937008,
        "ambience_sentiment": 0.7047244094488189,
        "price_sentiment": 0.8149606299212598,
        "overall_sentiment": 0.8248031496062992
    }
}
```

The base model agrees with GPT-4o on 82.7% of the labels overall, and does worst on ambience.

## Fine-tune the model

Turn each training row into a chat example, where the assistant message is the GPT-4o label. Then upload the file to Together.

````python theme={null}
def format_response(row):
    labels = {aspect: row[f"{aspect}_gt"] for aspect in ASPECTS}
    return f"```json\n{json.dumps(labels, indent=4)}\n```"


with open("finetuning_dataset.jsonl", "w") as f:
    for row in train_dataset:
        example = {
            "messages": [
                {"role": "system", "content": PROMPT},
                {"role": "user", "content": f"The review is: {row['text']}"},
                {"role": "assistant", "content": format_response(row)},
            ]
        }
        f.write(json.dumps(example) + "\n")

client = Together()
file = client.files.upload("finetuning_dataset.jsonl")
````

Start a LoRA fine-tuning job.

```python theme={null}
fine_tune_response = client.fine_tuning.create(
    training_file=file.id,
    model="meta-llama/Meta-Llama-3.1-8B-Instruct-Reference",
    n_epochs=3,
    suffix="aspect-based-sentiment-analysis-lora",
    lora=True,
    lora_r=64,
    wandb_api_key=os.environ.get("WANDB_API_KEY"),
)
```

Check the job until it finishes. Replace `ft-xyz` with your job ID.

```bash theme={null}
together fine-tuning list-events ft-xyz
```

<Note />

## Evaluate the fine-tuned model

When the job is done, find your model ID on the [Together models page](https://api.together.xyz/models) and run it on the test set.

```python theme={null}
ft_output = AspectBasedSentimentCurator(
    # Replace this with the ID of your fine-tuned model.
    model_name="together_ai/<your-together-username>/Meta-Llama-3.1-8B-Instruct-Reference-aspect-based-sentiment-analysis-lora-<job-suffix>",
    backend="litellm",
    generation_params={"temperature": 0.0},
    backend_params={"max_tokens_per_minute": 100_000_000},
)(test_dataset).dataset

ft_eval = evaluate_sentiment(ft_output)
print(json.dumps(ft_eval, indent=4))
```

```json theme={null}
{
    "overall_accuracy": 0.917716535433071,
    "aspect_accuracies": {
        "food_sentiment": 0.9035433070866141,
        "service_sentiment": 0.9409448818897638,
        "ambience_sentiment": 0.8858267716535433,
        "price_sentiment": 0.8937007874015748,
        "overall_sentiment": 0.9645669291338582
    }
}
```

## Compare the results

```python theme={null}
import pandas as pd

comparison_df = pd.DataFrame({
    "Metric": ["Overall Accuracy"] + [k.replace("_", " ").title() for k in base_eval["aspect_accuracies"]],
    "Base Model": [base_eval["overall_accuracy"]] + list(base_eval["aspect_accuracies"].values()),
    "Fine-tuned Model": [ft_eval["overall_accuracy"]] + list(ft_eval["aspect_accuracies"].values()),
})
pct = (comparison_df["Fine-tuned Model"] - comparison_df["Base Model"]) / comparison_df["Base Model"] * 100
comparison_df["Improvement"] = pct.apply(lambda x: f"{x:.2f}%")
print(comparison_df.round(3))
```

| Metric | Base model | Fine-tuned model | Improvement |
| - | - | - | - |
| Overall accuracy | 0.827 | 0.918 | 10.95% |
| Food sentiment | 0.866 | 0.904 | 4.32% |
| Service sentiment | 0.925 | 0.941 | 1.70% |
| Ambience sentiment | 0.705 | 0.886 | 25.70% |
| Price sentiment | 0.815 | 0.894 | 9.66% |
| Overall sentiment | 0.825 | 0.965 | 16.95% |

The fine-tuned model agrees with GPT-4o more often than the base model does, overall and for every aspect. It is also much cheaper to run. When this example was written, the 8B model cost $0.18 per million tokens on Together, and GPT-4o cost $2.50 per million input tokens. To get closer to GPT-4o, you can label a larger dataset and tune the training settings.
