DataGenerationConfig

Generate synthetic data using a model for finetuning an LLM.

KEY TYPE Description
descriptionCol str Name of the description column.
frequencyPenalty float Penalty for frequency of token appearance.
fewshotExamples int Number of fewshot examples used to prompt the model.
temperature float Sampling temperature for the model.
examplesPerTarget int Number of examples per target.
oversample bool Whether to oversample the data.
documentationCharLimit int Character limit for documentation.
model str Model to use for data generation.
idCol str Name of the identifier column.
generationInstructions str Instructions for the data generation model.
subsetSize Optional[int] Size of the subset to use for generation.
verifyResponse bool Whether to verify the response.
tokenBudget int Token budget for generation.
concurrency int Number of concurrent processes.
seed Optional[int] Seed for random number generation.
completionCol str Name of the output completion column.
promptCol str Name of the input prompt column.