DOpus 13.24.1 treats iso-2022-jp as shift_jis, but should not

Directory Opus 13.24.1 (Beta) said:

StringTools Encode/Decode methods now recognize the "WebCharset" keywords from the MIME types list in the registry (e.g. "shift_jis" is the web charset name for "iso-2022-jp").

But unfortunately, shift_jis and iso-2022-jp are related, but not the same encoding.

I checked this DOpus treats iso-2022-jp as shift_jis, by using Ad-Hoc Script Editor and python.

JScript on DOpus:

var stringTools = DOpus.Create().StringTools();
DOpus.Output(stringTools.Encode(stringTools.Encode("ใƒ†ใ‚นใƒˆ", "iso-2022-jp"), "base64"))

got this base64:

g2WDWINn

and it can decode as shift_jis using python:

>>> import base64
>>> base64.decodebytes(b'g2WDWINn').decode('shift_jis')
'ใƒ†ใ‚นใƒˆ'

But iso-2022-jp is stateful (with escape-sequence) encoding:

>>> 'ใƒ†ใ‚นใƒˆ'.encode('iso-2022-jp')
b'\x1b$B%F%9%H\x1b(B'

I believe that BodyCharset is "recommended charset on whatever for that codepage", not "web alias name of that codepage".

My MIME Database cp932 information:

reg query HKCR\MIME\Database\Codepage\932

HKEY_CLASSES_ROOT\MIME\Database\Codepage\932
    BodyCharset    REG_SZ    iso-2022-jp
    Description    REG_SZ    @%SystemRoot%\system32\mlang.dll,-4647
    Encoding    REG_BINARY    01010000
    FixedWidthFont    REG_SZ    MS Gothic
    Level    REG_BINARY    07070000
    ProportionalFont    REG_SZ    MS PGothic
    WebCharset    REG_SZ    shift_jis

Personally, I'm using this table for a handy alias reference.

This table treats shift_jis and cp932 as not the same encoding, and they're a bit different technically, but you can ignore and treat cp932 as shift_jis if you can't handle them differently.

Opus maps the string you give it to a codepage via the registry. Based on that, both iso-2022-jp and shift_jis map to codepage 932. The string conversion is done using the underlying codepage so there isn't really anything else shift_jis could be mapped to. Maybe it's technically incorrect (I'm not an expert on this) but you'd need to talk to Microsoft about creating a new codepage to fix it.

Would you like to provide the official reference of this registry key?

I found Code Page Identifiers - Win32 apps | Microsoft Learn code page identifiers table, and I found that some of the BodyCharsets are defined as another code page.

Also, I enumerated BodyCharset values for all code pages and found that some BodyCharsets are repeated. So I concluded that you should not believe that BodyCharset is an alias for codepage.

codepage .NET Name BodyCharset remarks by me
708 ASMO-708 ASMO-708
720 DOS-720 DOS-720
852 ibm850 ibm852
862 DOS-862 DOS-862
866 cp866 cp866
874 windows-874 windows-874
932 shift_jis iso-2022-jp This codepage isn't compatible with iso-2022-jp
936 gb2312 gb2312
949 ks_c_5601-1987 euc-kr might be compatible, but euc-kr defined as 51949
950 big5 big5
1200 utf-16 unicode (idk; might be related)
1201 unicodeFFFE unicodeFFFE
1250 windows-1250 iso-8859-2 might be related, but iso-8859-2 defined as 28592
1251 windows-1251 koi8-r might be related, but koi8-r defined as 20866
1252 windows-1252 iso-8859-1 might be related, but iso-8859-1 defined as 28591
1253 windows-1253 iso-8859-7 might be related, but iso-8859-7 defined as 28597
1254 windows-1254 iso-8859-9 might be related, but iso-8859-9 defined as 28599
1255 windows-1255 iso-8859-8-i might be related, but iso-8859-8-i defined as 38598
1256 windows-1256 iso-8859-6 might be related, but iso-8859-6 defined as 28596
1257 windows-1257 iso-8859-4 might be related, but iso-8859-4 defined as 28594
1258 windows-1258 windows-1258
20866 koi8-r koi8-r This should be real koi8-r instead of 1251
21866 koi8-u koi8-ru (idk; might be related)
28592 iso-8859-2 iso-8859-2 This should be real iso-8859-2 instead of 1250
28593 iso-8859-3 iso-8859-3
28594 iso-8859-4 iso-8859-4 This should be real iso-8859-4 instead of 1257
28595 iso-8859-5 iso-8859-5
28596 iso-8859-6 iso-8859-6 This should be real iso-8859-6 instead of 1256
28597 iso-8859-7 iso-8859-7 This should be real iso-8859-7 instead of 1253
28598 iso-8859-8 iso-8859-8
38598 iso-8859-8-i iso-8859-8-i This should be real iso-8859-8-i instead of 1255
50000 x-user-defined (unknown)
50001 _autodetect_all (unknown)
50220 iso-2022-jp iso-2022-jp This should be real iso-2022-jp instead of 932
50221 _iso-2022-jp$ESC (idk; related to 50220)
50222 _iso-2022-jp$SIO (idk; related to 50220)
50225 iso-2022-kr iso-2022-kr
50932 _autodetect (unknown)
50949 _autodetect_kr (unknown)
51932 euc-jp euc-jp
51949 euc-kr euc-kr I think that this shuld be real euc-kr instead of 949
52936 hz-gb-2312 hz-gb-2312
65000 utf-7 utf-7
65001 utf-8 utf-8

The best thing to do is to use the codepage directly. The lookup Opus does is provided as a convenience, at the end of the day the cp is what's used to do the conversion.

I have now found that you should use HKCR\MIME\Database\Charset\[CHARSET NAME] for codepage lookups. It has a unique codepage definition for each charset. But InternetEncoding seems to be the right value name for that instead of Codepage if that key has it.

I decided to keep using code page personally, but I don't want you to keep wrong aliasing. I want to prevent another script developer from relying on that. They might publish the script with a user-customizable charset ... for example, sending an email? ... with that. And the script developer will be confused by that user's bug report.

If I change it to enumerate the Charset key rather than the Codepage key we end up with this for the lookup table; does this seem right to you now?

name codepage
_autodetect 50932
_autodetect_all 50001
_autodetect_kr 50949
_iso-2022-jp$ESC 50221
_iso-2022-jp$SIO 50222
ANSI_X3.4-1968 1252
ANSI_X3.4-1986 1252
arabic 28596
ascii 1252
ASMO-708 708
Big5 950
chinese 936
CN-GB 936
cp1256 1256
cp367 1252
cp819 1252
cp852 852
cp866 866
csASCII 1252
csbig5 950
csEUCKR 949
csEUCPkdFmtJapanese 51932
csGB2312 936
csISO2022JP 50221
csISO2022KR 50225
csISO58GB231280 936
csISOLatin1 1252
csISOLatin2 28592
csISOLatin4 28594
csISOLatin5 1254
csISOLatinArabic 28596
csISOLatinCyrillic 28595
csISOLatinGreek 28597
csISOLatinHebrew 28598
csKOI8R 20866
csKSC56011987 949
csShiftJIS 932
csUnicode11UTF7 65000
csWindows31J 932
cyrillic 28595
DOS-720 720
DOS-862 862
DOS-874 874
ECMA-114 28596
ECMA-118 28597
ELOT_928 28597
euc-jp 51932
euc-kr 949
Extended_UNIX_Code_Packed_Format_for_Japanese 51932
GB2312 936
GB_2312-80 936
GBK 936
greek 28597
greek8 28597
hebrew 28598
hz-gb-2312 52936
IBM367 1252
ibm819 1252
ibm852 852
ibm866 866
iso-2022-jp 50220
iso-2022-kr 50225
iso-8859-1 1252
iso-8859-11 874
iso-8859-2 28592
iso-8859-3 28593
iso-8859-4 28594
iso-8859-5 28595
iso-8859-6 28596
iso-8859-7 28597
iso-8859-8 28598
ISO-8859-8 Visual 28598
iso-8859-8-i 38598
iso-8859-9 1254
iso-ir-100 1252
iso-ir-101 28592
iso-ir-110 28594
iso-ir-111 28594
iso-ir-126 28597
iso-ir-127 28596
iso-ir-138 28598
iso-ir-144 28595
iso-ir-148 1254
iso-ir-149 949
iso-ir-58 936
iso-ir-6 1252
ISO646-US 1252
iso8859-1 1252
iso8859-2 28592
ISO_646.irv:1991 1252
iso_8859-1 1252
iso_8859-1:1987 1252
iso_8859-2 28592
iso_8859-2:1987 28592
ISO_8859-4 28594
ISO_8859-4:1988 28594
ISO_8859-5 28595
ISO_8859-5:1988 28595
ISO_8859-6 28596
ISO_8859-6:1987 28596
ISO_8859-7 28597
ISO_8859-7:1987 28597
ISO_8859-8 28598
ISO_8859-8:1988 28598
ISO_8859-9 1254
ISO_8859-9:1989 1254
koi 20866
koi8-r 20866
koi8-ru 21866
korean 949
ks_c_5601 949
ks_c_5601-1987 949
ks_c_5601-1989 949
KSC5601 949
KSC_5601 949
l1 1252
l2 28592
l4 28594
l5 1254
latin1 1252
latin2 28592
latin4 28594
latin5 1254
logical 1255
ms_Kanji 932
shift-jis 932
shift_jis 932
unicode 1200
unicode-1-1-utf-7 65000
unicode-1-1-utf-8 65001
unicode-2-0-utf-8 65001
unicodeFFFE 1201
us 1252
us-ascii 1252
utf-7 65000
utf-8 65001
visual 28598
windows-1250 1250
windows-1251 1251
windows-1252 1252
windows-1253 1253
Windows-1254 1254
windows-1255 1255
windows-1256 1256
windows-1257 1257
windows-1258 1258
windows-874 874
x-ansi 1252
x-cp1250 1250
x-cp1251 1251
x-euc 51932
x-euc-jp 51932
x-ms-cp932 932
x-sjis 932
x-unicode-2-0-utf-7 65000
x-unicode-2-0-utf-8 65001
x-user-defined 50000
x-x-big5 950

Thanks! I checked with the Code Page Identifiers table (above, microsoft.com), no known bad aliases. I believe this works well.

Great, we'll change to that method in the next beta. Thanks for your help with this!