LLARS: Enabling Domain Expert & Developer Collaboration for LLM Prompting, Generation and Evaluation
Researchers have introduced LLARS (LLM Assisted Research System), an open-source platform designed to bridge the collaboration gap between domain experts and developers in building Large Language Model (LLM) systems. The platform integrates three core modules into a seamless end-to-end pipeline: Collaborative Prompt Engineering, which supports real-time co-authoring with version control; Batch Generation, allowing configurable output production across various prompts, models, and datasets with cost controls; and Hybrid Evaluation, where human and LLM evaluators jointly assess outputs using diverse methods and live agreement metrics. This system enables users to identify optimal model-prompt combinations through provenance analysis. Validated through interviews with domain experts and developers in online counseling, LLARS was reported to be intuitive and time-saving by centralizing workflows. The platform facilitates interdisciplinary cooperation by automatically making new prompts available for batch generation and converting completed batches into evaluation scenarios instantly. This development aims to streamline the creation and assessment of LLM-based applications, enhancing efficiency in research and development processes.
Wire timeline
LLARS: Enabling Domain Expert & Developer Collaboration for LLM Prompting, Generation and Evaluation
Researchers have introduced LLARS (LLM Assisted Research System), an open-source platform designed to bridge the collaboration gap between domain experts and developers in building Large Language Model (LLM) systems. The platform integrates three core modules into a seamless end-to-end pipeline: Collaborative Prompt Engineering, which supports real-time co-authoring with version control; Batch Generation, allowing configurable output production across various prompts, models, and datasets with cost controls; and Hybrid Evaluation, where human and LLM evaluators jointly assess outputs using diverse methods and live agreement metrics. This system enables users to identify optimal model-prompt combinations through provenance analysis. Validated through interviews with domain experts and developers in online counseling, LLARS was reported to be intuitive and time-saving by centralizing workflows. The platform facilitates interdisciplinary cooperation by automatically making new prompts available for batch generation and converting completed batches into evaluation scenarios instantly. This development aims to streamline the creation and assessment of LLM-based applications, enhancing efficiency in research and development processes.
cs.AI updates on arXiv.org