Miniblog

Short thoughts

  • 一直知道台灣花磚博物館,卻沒機會拜訪。今天看到館方推行的花磚日曆集資專案,頗為有趣。介紹說明有兩百餘頁日曆附著QR碼,可以徑行前往花磚所在地。

    等待之餘,發現國家文化記憶庫藏有館方上傳的花磚照片,可以望梅止渴。

    Posted 4 days ago
  • 前些日子,Pan's Labyrinth的導演Guillermo del Toro做了一個AMA,其中一個回答提及了帶給他靈感的兩本書:

    When I was writing the movie*, I locked myself in a little room in a hotel in Spain. And I brought six or seven Victorian and Edwardian volumes of fairy lore and fairy tales. One was a couple of volumes from Andrew Lang’s colored fairy books. Another one was, The Science of Fairy Tales, which is a beautiful book, and then some of the taxonomy for the fairies in the highlands [...]

    其一為維多利亞時代紅極一時的Andrew Lang彩色童話集系列*(見圖),其二為 The Science of Fairy Tales。*

    一直對童話深感興趣,期許這些書也能帶給我些靈感。

    即Pan's Labyrinth。 Archive.orgArchive.org
    Posted 5 days ago · Updated 4 days ago
  • I previously mentioned that I was using some websites to discover all kinds of blogs. Here are some I found all along the way:

    Posted 5 days ago
  • My token budget is starting to burn fast. Here are the token optimizations I've done to reduce the cost for OpenCode.

    • Install RTK.
      • Reduce shell output. See here for the supported commands.
    • Install Codegraph.
      • Remember to codegraph init new repository !
      • Make sure that the MCP server is live.
    • Install OpenSlimedit.
      • The code source is very tiny. It basically compress built-in tool descriptions that are frequently used, and perform additional little optimisations here and there.
    • Use cheaper models for subagents.

    Excluded plugins

    • Dynamic Context Pruning Plugin : Not particularily convinced that it would reduce cost. Especially when I read their section Impact on Prompt Caching that kinda confirm my skepticism :

      Trade-off: Pruning reduces context size but can increase cache misses. The cost balance depends on your conversation, compression frequency, and provider pricing.

      Cache misses were exactly a significant part of my tokens cost. No point to risk it only to figure out what "intelligently" managed conversation ever mean.

    • OpenCode Snip : Used to reduce the output of shell commands. Most certainly redunding with RTK.

    Posted 2 wk. ago · Updated last wk.
  • 方才在Lichess戰勝了同一位對手,看了一下記錄,不想這位兄臺時隔一週居然又掉入了一模一樣的陷阱*。即

    1. d4 Nc6 2. Bf4 h5!? 3. e3? e5!

    對手黑象就插翅難飛了。

    以上局面來自極罕*的Mikėnas防禦(1. d4 Nc6)。實際上,上述對付倫敦系統的陷阱可以廣義化成:

    1. d4 * 2. Bf4 h5!? 3. e3? e5!

    任何對d4的主流回覆都行得通。

    這陷阱適合用以懲戒某些倫敦系統玩家無視對手下法的壞毛病。

    後記:過幾天又遇上了,還是中招了。都懷疑他是不是bot。 < 1%
    Posted 2 wk. ago · Updated last wk.
  • 除了主要使用的LLM之外,有時會用到Deepseek的API。

    上個月,該公司公佈了要改採尖離峰價,價格差達兩倍之多。尖峰時間,簡單而言就是北京上班時間(不包括中午休息及大陸地區假日)。

    對東亞時區的使用者來說,Deepseek或許不再有低價的吸引力。對歐洲就絲毫無任何改變,因為離峰時間對應的是下午至午夜。

    某位兄臺為此做了一面DeepSeek尖離峰時段價格時鐘,供大家參考。

    Posted 2 wk. ago
  • 這時代還有些大公司仍用著千篇一律的拖曳式拼圖Captcha。

    VLM一下就能低成本寫出一個CV-based的通解*,也不消仰賴諸如2Captcha的人類API了。

    這麼多年了,Captcha依舊阻礙人類,卻讓程式通行無阻,何嘗不是一種逆圖靈測試?

    亦即是說成本比放置Captcha的公司還低(零)。換言之,公司在花錢降低UX。
    Posted 2 wk. ago · Updated 2 wk. ago
  • 長時間半夜看螢幕實在傷眼,想著應該用redshift護眼,不想又跟Wayland不合*。只好使用Ubuntu內建的Night Light功能。結果不知怎的,只有兩面螢幕有效果。原來DisplayLink不支援校正,九年前的feature request也視若無睹。最終只好自己調整剩餘螢幕的亮度。

    對Wayland的厭惡之情日漸增長。
    Posted 2 wk. ago · Updated 2 wk. ago
  • I absolutely love Richard Stallman's glossary and anti-glossary. Though I do not necessarily share all of the pointviews of our church leader, St. IGNUcius, I might actually start to use some of the terms listed there.

    BIBO (Bias In, Bias Out) is a brilliant one. Sounds like a variant of "garbage in, garbage out" (GIGO?), both occuring in AI system training, but also FIFO (First In, First Out) in data structure manipulatiom.

    Posted 2 wk. ago · Updated 2 wk. ago
  • 某GAFAM旗下,擁有超過一億用戶的子公司,受著業界龍頭Datadome的保護,不想可以輕易繞過,任人隨意建立帳號,實在不可思議。這就是所謂的「紙老虎」罷。

    Posted 2 wk. ago
  • 作為一名專下棄兵開局*的gambiteer*,可參考的開局書籍常常寥寥無幾。

    Tennison棄兵就是這麼樣的開局。

    斯堪的納維亞防禦(1. e4 d5)一直是我的剋星,嘗試過主變例(2. exd5)及較差但個人勝率較高的前進變例(2. e5)。最後相中了這個Tennison棄兵納入我的開局錄*,並一連幾周的晚上在棋社與棋友嘗試各種變例始有心得。後在今年Noisiel的比賽,受到Martian棄兵的啟發,局中找到引擎合意的棄馬妙手,不表。

    Tennison棄兵的最後一本較正式的書籍應該是Uwe Bekemann的Better late than never - The Tennison Gambit(2015),其簡介還記載了開局典故*:

    This is one of these old stories from chess history which can neither be confirmed nor ultimately refuted. According to the story, and so far there is no doubt about its authenticity, Otto Tennison was a Danish player who lived in the 19th century. He used to open all his games categorically with the strongest first move 1.e4 until one day he became so fed up with all the elaborate variations his opponents threw onto the board without even thinking that he said to himself, "Enough is enough! From now on I will choose a completely different approach.“ Thus spoke Otto Tennison, and at the next opportunity when he played with the white pieces, he threw 1.Nf3 onto the board without even thinking. Alas, when his opponent answered 1...d5, he was overcome by doubts about what he had done. He sat and stared at the board as if in deep meditation, but suddenly, like struck by lightning, he took his king pawn, pushed it to e4 with great decisiveness and uttered – more to himself than to anybody else, "Better late than never!“ The Tennison Gambit is one of these openings which the chess world has almost entirely neglected up to now. The initial move order 1.Nf3 d5 2.e4 (or less often 1.e4 d5 2.Nf3) leads to positions which are new territory for most players. As a surprise weapon it is the ideal approach to lure your opponent into theoretical no man's land almost from move one.

    即「gambit」。 當如何迻譯?「棄子俠」?林紓味有點重。 「Opening repertoire」的個人暫譯。 通常拿它來對付斯堪的納維亞防禦 1. e4 d5 2. Nf3,所以「Better late than never」不適用於我。
    Posted 3 wk. ago · Updated 2 wk. ago
  • 通過國立故宮博物院開放的「清代檔案檢索系統」讀清代奏折收穫頗豐。清代書面語尚屬簡單易懂,學習新詞尤為方便。作為日常讀物,內容也有幾分意趣。其中臣子的奴性,或有不及當代「牛馬」。

    閱讀奏折觀察道以下幾點:

    1. 數字大寫:用於計量、日期。
    2. 格式:
      1. 無條件換行:每逢「皇上」、「聖~」、「恩」等必換行。
    3. 避君諱:因為內容性質的關係,目前還未觀察到,不然應該可以預期康熙朝奏折有諸如「玄機」 --> 「元機」,可用於確定年號的避諱。
    4. 異體字:「熱」作「𤍠」、「旨」作「㫖」等。

    2021年開放的平台,OCR不甚好(pdf內文有識別錯誤,網站上的硃批則有漏字)。從去歲旁聽的線上研討會得知以現在的技術來看,要提高準確率並非難事。

    Posted 3 wk. ago · Updated last wk.
  • The National Palace Museum Collection (臺北故宮) has made open data a quite substantial amount of its collection (108k items) consisting of high quality images, metadata and description with terminological precision. I see two obvious immediate direct applications :

    1. Building a Chinese Iconography Thesaurus. The description already contain pattern names like 勾連雲雷紋 We would then need to apply image transformation and maybe use VLM to extract and highlight the motif.
    2. Buidling a specialized glossary. Would be handy for Chinese literary translators and Chinese lexicography.

    The existing Chinese Iconography Thesaurus is a nice effort, but seems to cover only a small amount (~13k items) of the collections available from the museums over the world.

    Data of similar quality can be obtained from auction houses as well, like Christie's.

    Posted 3 wk. ago · Updated 3 wk. ago
  • I'm using anyblog.cc to randomly find blogs that could inspire me, being the content, the writing style, or the design. It just lead me to the webpage of Gitlab's Co-founder, Sid Sijbrandij. He got cancer, and surprisingly made open source his personal genomics and imaging data (Gitlab spirit !) A dozen of biotech companies has been found during the process to help others. What a madman !

    Posted 3 wk. ago · Updated 3 wk. ago