Abstract:
Artificial intelligence generated content (AIGC) has seriously affected information authenticity and reliability, leading to various technical and social problems such as data pollution, property ownership, and credibility crisis. Existing machine-generated text detection methods are primarily designed for specific domains and suffer from relatively low detection accuracy, making them even less effective when applied to cross-domain data such as sensitive, private, or small-sample data. To address this problem, a high available cross-domain machine-generated text detection method was proposed. This method first selected the class-center samples in any domain to train a domain-specific encoder, thereby leveraging domain features enhance boundary distinguishability. Then, an orthogonal loss function was constructed to train a domain-general encoder with the domain-specific encoder, reinforcing the general-feature of machine-generated text to support the detection across multiple domains. Experimental results on real-world data show that the detection model trained on a single domain can obtain high detection accuracy in other domains without fine-tuning, highlighting its broad applications and strong practicality.