I have seen technology sharing and company introductions related to llm eval more than once this year. The real effect evaluation feedback of LLM will serve as the direction of the company's subsequent optimization, exposing the true situation of existing AI capabilities. I used to simply think of it as just doing some evaluations. In fact, llm evaluation is getting more and more attention, especially on the b-side, which pursues stable, reliable and high-quality data. Anthropic's acquisition of humanloop on the one hand proves that llm eval is indeed very important, and on the other hand it hints that Anthropic really wants to make money. The money is on the b-side, because 1.6 billion of their previous revenue of more than 5 billion US dollars came from cursor. Obviously this toc model is unhealthy. It can only last for three years after getting a few big orders after developing strength on the b-side. Another point, it seems that I have never heard of specializing in llm eavl in China. It may be a good opportunity.


Visual summary
Article argument map
Generated from the post's content graph
Mermaid outline
flowchart LR
format-long_post["longform note"]
topic-evaluation["evaluation"]
capability-evaluation["evaluation"]
format-long_post -->|covers| topic-evaluation
format-long_post -->|signals| capability-evaluationVisual structure
Essay structure map
Built from summary and key paragraph positions
Mermaid outline
flowchart LR
thesis["I have seen technology sharing and company introductions related to llm eval more than once this year. The..."]
signal["I have seen technology sharing and company introductions related to llm eval more than once this year. The..."]
operator["I have seen technology sharing and company introductions related to llm eval more than once this year. The..."]
implication["I have seen technology sharing and company introductions related to llm eval more than once this year. The..."]
thesis -->|frames| signal
signal -->|develops| operator
operator -->|lands in| implication