ambience_sentiment, price_sentiment, and overall_sentiment.
Setup
Label the data
The prompt asks the model to rate five aspects of a review. This example does not use structured output, because the sameLLM class also runs the small base model later, and many small models do not support structured output. The prompt asks for JSON in a code block instead.
Split the data
Rename the label columns with a_gt suffix, for ground truth. Then use 90% of the rows for training and 10% for testing.
Evaluate the base model
This function compares the model output with the GPT-4o labels and returns the accuracy for each aspect.Fine-tune the model
Turn each training row into a chat example, where the assistant message is the GPT-4o label. Then upload the file to Together.ft-xyz with your job ID.
Evaluate the fine-tuned model
When the job is done, find your model ID on the Together models page and run it on the test set.Compare the results
The fine-tuned model agrees with GPT-4o more often than the base model does, overall and for every aspect. It is also much cheaper to run. When this example was written, the 8B model cost 2.50 per million input tokens. To get closer to GPT-4o, you can label a larger dataset and tune the training settings.