Multi-keyword secure search of encrypted cloud data based on semantics

Xuejian Wang, Baodong Qin · 2024

Most traditional searchable encryption schemes extract features from each document to construct a vector representation of the document, such as the TF-IDF model. These vectors not only have high dimensionality and complex operations, but also ignore the semantic information of the documents themselves. When calculating the importance of a document, the TF-IDF model usually needs to take into account word frequency and inverse document frequency. This means that this information needs to be stored in the encrypted index, thus potentially exposing part of the document’s information, especially for words that have a high word frequency or occur in most of the documents. SBERT is a model for generating sentence embeddings, which adapts and extends the BERT model to generate high-quality semantic representations at the sentence level. Using the SBERT model, vectors containing semantic information of documents can be trained to realize semantic-aware multi-keyword search. By training this model, more accurate search results can be obtained in document retrieval.

Read the paper · More papers on PaperTik