迷你部落格

短文隨想

篩選: open-data 全部標籤

  • 通過國立故宮博物院開放的「清代檔案檢索系統」讀清代奏折收穫頗豐。清代書面語尚屬簡單易懂,學習新詞尤為方便。作為日常讀物,內容也有幾分意趣。其中臣子的奴性,或有不及當代「牛馬」。

    閱讀奏折觀察道以下幾點:

    1. 數字大寫:用於計量、日期。
    2. 格式:
      1. 無條件換行:每逢「皇上」、「聖~」、「恩」等必換行。
    3. 避君諱:因為內容性質的關係,目前還未觀察到,不然應該可以預期康熙朝奏折有諸如「玄機」 --> 「元機」,可用於確定年號的避諱。
    4. 異體字:「熱」作「𤍠」、「旨」作「㫖」等。

    2021年開放的平台,OCR不甚好(pdf內文有識別錯誤,網站上的硃批則有漏字)。從去歲旁聽的線上研討會得知以現在的技術來看,要提高準確率並非難事。

    發布於 3 週前 · 更新於 上週
  • The National Palace Museum Collection (臺北故宮) has made open data a quite substantial amount of its collection (108k items) consisting of high quality images, metadata and description with terminological precision. I see two obvious immediate direct applications :

    1. Building a Chinese Iconography Thesaurus. The description already contain pattern names like 勾連雲雷紋 We would then need to apply image transformation and maybe use VLM to extract and highlight the motif.
    2. Buidling a specialized glossary. Would be handy for Chinese literary translators and Chinese lexicography.

    The existing Chinese Iconography Thesaurus is a nice effort, but seems to cover only a small amount (~13k items) of the collections available from the museums over the world.

    Data of similar quality can be obtained from auction houses as well, like Christie's.

    發布於 3 週前 · 更新於 3 週前
  • I'm using anyblog.cc to randomly find blogs that could inspire me, being the content, the writing style, or the design. It just lead me to the webpage of Gitlab's Co-founder, Sid Sijbrandij. He got cancer, and surprisingly made open source his personal genomics and imaging data (Gitlab spirit !) A dozen of biotech companies has been found during the process to help others. What a madman !

    發布於 3 週前 · 更新於 3 週前