C/C++ - 리눅스 환경에서 u16string 문자열을 출력하는 방법

우선, u16string은 직접적으로 cout에 출력 지원이 안 됩니다.

std::u16string text = u"test";
cout << text << endl;

위와 같이 하면 이런 식으로 컴파일 오류가 발생합니다.

message : cannot convert ‘text’ (type ‘std::u16string’ {aka ‘std::__cxx11::basic_string<char16_t>’}) to type ‘const unsigned char*’

보통, 이런 경우 c_str() 함수의 결과를 출력하는데, u16string 계열은 이것마저도 단순히 해당 문자열의 주솟값을 출력할 뿐입니다.

cout << text.c_str() << endl; // 출력 결과: 0x7fffffffe330

답답하군요. ^^ 그럼, 어차피 16비트니까 wchar_t로 형변환하면 되지 않을까요? 그런데 실제로 해보면 정상적인 출력이 안 나옵니다.

const wchar_t* result = (wchar_t*)text.c_str();
wcout << result << endl; // 출력 결과: ts U  U

text의 메모리 표현이 "74 00 65 00 73 00 74 00 00"으로 나오는데, 왜 wchar_t로 정상적으로 받을 수 없는 걸까요? 그것은 리눅스에서 wchar_t 타입이 윈도우처럼 2바이트가 아닌 4바이트이기 때문입니다. 그로 인해 "74 00 65 00"이 한 글자를 표현하게 되고 결국 4바이트 유니코드에 해당하는 문자를 나타내 의도치 않은 결과가 나옵니다.

그래서 검색해 보면, u16string을 utf8로 변환 후 출력하라는 글들이 나옵니다.

how to print u32string and u16string to the console in c++
; https://stackoverflow.com/questions/45874857/how-to-print-u32string-and-u16string-to-the-console-in-c

std::wstring_convert<std::codecvt_utf8<char16_t>, char16_t> converter;
std::cout << converter.to_bytes(text) << std::endl; // 출력 결과: test

한 가지 유의할 사항은, converter.to_bytes 함수가 변환할 수 없는 문자를 포함하고 있으면 std::range_error 예외를 발생한다는 점입니다. 대개의 경우 정상적으로 초기화하지 않은 버퍼가 입력으로 들어간 경우에 발생할 텐데요, 이 외에도 Surrogate Pair를 지원하지 않아,

유니코드의 Surrogate Pair, Supplementary Characters가 뭘까요?
; https://www.sysnet.pe.kr/2/0/1710

정상적인 UTF-16 인코딩 문자열이라도 예외가 발생할 수 있음에 주의해야 합니다.

std::u16string text = u"\xd800\xdc00"; // U+10000에 대한 Surrogate Pair

std::wstring_convert<std::codecvt_utf8<char16_t>, char16_t> converter;
std::cout << converter.to_bytes(text) << std::endl; // 예외 발생

/*
terminate called after throwing an instance of 'std::range_error'
  what():  wstring_convert::to_bytes
*/

그래서, 가능한 try/catch로 to_bytes를 보호하는 것이 좋습니다.

try 
{
    std::u16string text = u"\xd800\xdc00";

    std::wstring_convert<std::codecvt_utf8<char16_t>, char16_t> converter;
    std::cout << converter.to_bytes(text) << std::endl; // 예외 발생
}
catch (std::range_error& e) {
    // 예외 처리
}

사실 이에 대해서는 문서에 명시돼 있긴 합니다.

codecvt_utf8
; https://learn.microsoft.com/en-us/cpp/standard-library/codecvt-utf8-class

Represents a locale facet that converts between wide characters encoded as UCS-2 or UCS-4, and a byte stream encoded as UTF-8.

보는 바와 같이 UTF-16이 아닌 UCS-2를 지원할 뿐입니다. 그런데, u16string의 value_type이 char16_t이고, char16_t가 UTF-16 인코딩을 받는 것을 보면,

char16_t
; https://en.cppreference.com/w/c/string/multibyte/char16_t

뭔가 잘 안 맞는 부분이 있는 듯합니다. 단지, UCS-4는 지원하기 때문에 다음과 같이 변환할 수는 있습니다.

// How to decode surrogate characters encoded as UTF8?
// ; https://stackoverflow.com/questions/38293373/how-to-decode-surrogate-characters-encoded-as-utf8

const wchar_t* text = L"\U00010000";
std::wstring_convert<std::codecvt_utf8_utf16<wchar_t>> converter;

std::string utf8str = converter.to_bytes(text); // F0 90 80 80

결국 surrogate-pair에 해당하는 d800, dc00 바이트 스트림을 utf-8로 변환하는 방법은...? 직접 그에 대한 처리를 해야 합니다. (잘 찾아보면 누군가 만들어 둔 것이 있을지도... ^^)

[이 글에 대해서 여러분들과 의견을 공유하고 싶습니다. 틀리거나 미흡한 부분 또는 의문 사항이 있으시면 언제든 댓글 남겨주십시오.]

[최초 등록일: 10/11/2022]
[최종 수정일: 8/30/2024]

이 저작물은 크리에이티브 커먼즈 코리아 저작자표시-비영리-변경금지 2.0 대한민국 라이센스에 따라 이용하실 수 있습니다.

by SeongTae Jeong, mailto:techsharer at outlook.com

No	Writer	Date	Cnt.	Title	File(s)
13645	정성태	6/13/2024	9831	개발 환경 구성: 711. Visual Studio로 개발 시 기본 등록하는 dev tag 이미지로 Docker Desktop k8s에서 실행하는 방법
13644	정성태	6/12/2024	11130	닷넷: 2265. C# - System.Text.Json의 기본적인 (한글 등에서의) escape 처리 [1]
13643	정성태	6/12/2024	9962	오류 유형: 907. MySqlConnector 사용 시 System.IO.FileLoadException 오류
13642	정성태	6/11/2024	9624	스크립트: 65. 파이썬 - asgi 버전(2, 3)에 따라 달라지는 uvicorn 호스팅
13641	정성태	6/11/2024	10903	Linux: 71. Ubuntu 20.04를 22.04로 업데이트
13640	정성태	6/10/2024	11315	Phone: 21. C# MAUI - Android 환경에서의 파일 다운로드(DownloadManager)
13639	정성태	6/8/2024	10685	오류 유형: 906. C# MAUI - Android Emulator에서 "Waiting For Debugger"로 무한 대기
13638	정성태	6/8/2024	10826	오류 유형: 905. C# MAUI - 추가한 layout XML 파일이 Resource.Layout 멤버로 나오지 않는 문제
13637	정성태	6/6/2024	9779	Phone: 20. C# MAUI - 유튜브 동영상을 MediaElement로 재생하는 방법
13636	정성태	5/30/2024	9358	닷넷: 2264. C# - 형식 인자로 인터페이스를 갖는 제네릭 타입으로의 형변환	1
13635	정성태	5/29/2024	11379	Phone: 19. C# MAUI - 안드로이드 "Share" 대상으로 등록하는 방법
13634	정성태	5/24/2024	12023	Phone: 18. C# MAUI - 안드로이드 플랫폼에서의 Activity 제어 [1]
13633	정성태	5/22/2024	11299	스크립트: 64. 파이썬 - ASGI를 만족하는 최소한의 구현 코드
13632	정성태	5/20/2024	10033	Phone: 17. C# MAUI - Android 내에 Web 서비스 호스팅
13631	정성태	5/19/2024	11143	Phone: 16. C# MAUI - /Download 등의 공용 디렉터리에 접근하는 방법 [1]
13630	정성태	5/19/2024	10027	닷넷: 2263. C# - Thread가 Task보다 더 빠르다는 어떤 예제(?)
13629	정성태	5/18/2024	10558	개발 환경 구성: 710. Android - adb.exe를 이용한 파일 전송
13628	정성태	5/17/2024	9807	개발 환경 구성: 709. Windows - WHPX(Windows Hypervisor Platform)를 이용한 Android Emulator 가속
13627	정성태	5/17/2024	9980	오류 유형: 904. 파이썬 - UnicodeEncodeError: 'ascii' codec can't encode character '...' in position ...: ordinal not in range(128)
13626	정성태	5/15/2024	11349	Phone: 15. C# MAUI - MediaElement Source 경로 지정 방법	1
13625	정성태	5/14/2024	10391	닷넷: 2262. C# - Exception Filter 조건(when)을 갖는 catch 절의 IL 구조
13624	정성태	5/12/2024	10225	Phone: 14. C# - MAUI에서 MediaElement 사용	1
13623	정성태	5/11/2024	9733	닷넷: 2261. C# - 구글 OAuth의 JWT (JSON Web Tokens) 해석	1
13622	정성태	5/10/2024	11886	닷넷: 2260. C# - Google 로그인 연동 (ASP.NET 예제)	1
13621	정성태	5/10/2024	11175	오류 유형: 903. IISExpress - Failed to register URL "..." for site "..." application "/". Error description: Cannot create a file when that file already exists. (0x800700b7)
13620	정성태	5/9/2024	10232	VS.NET IDE: 190. Visual Studio가 node.exe를 경유해 Edge.exe를 띄우는 경우

Writer

Date

Cnt.

Title

File(s)

13645

정성태

6/13/2024

9831

개발 환경 구성: 711. Visual Studio로 개발 시 기본 등록하는 dev tag 이미지로 Docker Desktop k8s에서 실행하는 방법

13644

정성태