I'm new to PostgreSQL, and from the looks of it, it's a great database,
and I'll be using more of it in the future.
I had a quick question if anyone could clear this up. The documentation
for PostgreSQL (version 7.1, the version this server is using) says that
it supports multibyte character encodings like Unicode (which implies
UTF-16 encoding). Later on, the same page says that Unicode is
represented using UTF-8 encoding. UTF-8 is the 8-bit version of Unicode.
The multibyte version of Unicode is UTF-16.
So, which is it? If I create a database using Unicode as the encoding,
will the encoding be UTF-8 (singlebyte) or UTF-16 (multibyte)?
Thanks!
Rich
---------------------------(end of broadcast)---------------------------
TIP 9: the planner will ignore your desire to choose an index scan if your
joining column's datatypes do not match 3 3451
At 8:39 PM -0400 9/16/04, Richard Connamacher wrote: I'm new to PostgreSQL, and from the looks of it, it's a great database, and I'll be using more of it in the future.
I had a quick question if anyone could clear this up. The documentation for PostgreSQL (version 7.1, the version this server is using) says that it supports multibyte character encodings like Unicode (which implies UTF-16 encoding).
Don't confuse Unicode, the 'character set' and rules for characters,
represented by a sequence of abstract 32 bit integers, with
UTF-[8|16|32] which is a way to encode those abstract integers into a
stream of bytes someplace.
Later on, the same page says that Unicode is represented using UTF-8 encoding. UTF-8 is the 8-bit version of Unicode. The multibyte version of Unicode is UTF-16.
So, which is it? If I create a database using Unicode as the encoding, will the encoding be UTF-8 (singlebyte) or UTF-16 (multibyte)?
Erm... UTF-8 *is* a multibyte encoding. Up to 6 bytes per code point,
if things get really degenerate. (And, last I checked, means you can
have up to 70 bytes for really degenerate characters, but my memory
might be off (could be 80))
UTF-8, UTF-16, and UTF-32 will all encode Unicode characters just fine.
--
Dan
--------------------------------------it's like this-------------------
Dan Sugalski even samurai da*@sidhe.org have teddy bears and even
teddy bears get drunk
---------------------------(end of broadcast)---------------------------
TIP 7: don't forget to increase your free space map settings
On Sep 17, 2004, at 9:39 AM, Richard Connamacher wrote: UTF-8 is the 8-bit version of Unicode. The multibyte version of Unicode is UTF-16.
UTF-8 encodes characters with varying numbers of bytes, not just 1 byte
per character. IIRC, it's anywhere from 1 to 5 bytes, actually.
PostgreSQL uses UTF-8.
If you can, upgrade. 7.1 is nearing prehistoric. :)
Michael Glaesemann
grzm myrealbox com
---------------------------(end of broadcast)---------------------------
TIP 4: Don't 'kill -9' the postmaster
=> show client_encoding ;
client_encoding
-----------------
UNICODE
(1 ligne)
=> select char_length('a' ), bit_length('a') ;
char_length | bit_length
-------------+------------
1 | 8
(1 ligne)
# that's an accented "e"
=> select char_length('é' ), bit_length('é') ; ;
char_length | bit_length
-------------+------------
1 | 16 <= two bytes
(1 ligne)
pg does not simply store utf-8 data, it also understands it if you set
your encoding correctly (ie. initdb to UNICODE and client_encoding too so
that data doesn't get mangled on the way to the db). It will refuse to eat
illegal UTF8 characters too.
Once you try unicode, all the codepage mess starts to look old...
On Thu, 16 Sep 2004 20:39:48 -0400, Richard Connamacher
<ri*****@indiei mage.com> wrote: I'm new to PostgreSQL, and from the looks of it, it's a great database, and I'll be using more of it in the future.
I had a quick question if anyone could clear this up. The documentation for PostgreSQL (version 7.1, the version this server is using) says that it supports multibyte character encodings like Unicode (which implies UTF-16 encoding). Later on, the same page says that Unicode is represented using UTF-8 encoding. UTF-8 is the 8-bit version of Unicode. The multibyte version of Unicode is UTF-16.
So, which is it? If I create a database using Unicode as the encoding, will the encoding be UTF-8 (singlebyte) or UTF-16 (multibyte)?
Thanks! Rich
---------------------------(end of broadcast)--------------------------- TIP 9: the planner will ignore your desire to choose an index scan if your joining column's datatypes do not match
---------------------------(end of broadcast)---------------------------
TIP 7: don't forget to increase your free space map settings This thread has been closed and replies have been disabled. Please start a new discussion. Similar topics |
by: lawrence |
last post by:
Someone on www.php.net suggested using a seems_utf8() method to test
text for UTF-8 character encoding but didn't specify how to write such
a method. Can anyone suggest a test that might work? Something that
maybe gives 90% confidence that a given block of text is or is not
UTF-8 encoded?
|
by: Alban Hertroys |
last post by:
Another python/psycopg question, for which the solution is probably
quite simple; I just don't know where to look.
I have a query that inserts data originating from an utf-8 encoded XML
file. And guess what, it contains utf-8 encoded characters...
Now my problem is that psycopg will only accept queries of type str, so
how do I get my utf-8 encoded data into the DB?
I can't do query.encode('ascii'), that would be similar to:
>>> x =...
|
by: Haines Brown |
last post by:
I'm having trouble finding the character entity for the French
abbreviation for "number" (capital N followed by a small supercript
o, period).
My references are not listing it. Where would I find an answer to this
question (don't find it in the W3C_char_entities document).
--
Haines Brown
brownh@hartford-hwp.com
|
by: jmgonet |
last post by:
Hello everybody,
I'm having troubles loading a Xml string encoded in UTF-8.
If I try this code:
------------------------------
XmlDocument doc=new XmlDocument();
String s="<?xml version=\"1.0\" encoding=\"utf-8\"
standalone=\"yes\"?><a>Schönbühl</a>";
doc.LoadXml(s);
doc.Save("d:\\temp\\test.xml");
|
by: archana |
last post by:
Hi all,
can someone tell me difference between unicode and utf 8 or utf 18 and
which one is supporting more character set.
whic i should use to support character ucs-2.
I want to use ucs-2 character in streamreader and streamwriter.
How unicode and utf chacters are stored.
| |
by: Jimmy Shaw |
last post by:
Hi everybody,
Is there any SIMPLE way to convert from UTF-16 to UTF-32? I may be
mixed up, but is it possible that all UTF-16 "code points" that are 16
bits long appear just the same in UTF-32, but with zero padding and
hence no real conversion is necessary?
If I am completely wrong and some intricate conversion operation needs
to take place, can anyone give me some primer on the subject?
|
by: sheldon.regular |
last post by:
I am new to unicode so please bear with my stupidity.
I am doing the following in a Python IDE called Wing with Python 23.
äöü
äöü
'\xc3\xa4\xc3\xb6\xc3\xbc'
u'\xe4\xf6\xfc'
u'\xe4\xf6\xfc'
äöü
|
by: shreshth.luthra |
last post by:
Hi All,
I am having a GUI which accepts a Unicode string and searches a given
set of xml files for that string.
Now, i have 2 XML files both of them saved in UTF-8 format, having
characters of different language.
Although both of them are having UTF-8 as BoM, but only first file is
having UTF-8 defined in XML declration at the top of the XML file as
|
by: Jed |
last post by:
I have a form that needs to handle international characters withing the UTF-8
character set. I have tried all the recommended strategies for getting utf-8
characters from form input to email message and I cannot get it to work. I
need to stay with classic asp for this.
Here are some things I tried:
'CDONTS
Call msg.SetLocaleIDs(65001)
|
by: Allan Ebdrup |
last post by:
I hava an ajax web application where i hvae problems with UTF-8 encoding oc
chineese chars.
My Ajax webapplication runs in a HTML page that is UTF-8 Encoded.
I copy and paste some chineese chars from another HTML page viewed in IE7,
that is also UTF-8 encoded (search for "china" on google.com). I paste the
chineese chars into a content editable div.
My Ajax webservice compiles an XML where the data from the content editable
div is...
|
by: marktang |
last post by:
ONU (Optical Network Unit) is one of the key components for providing high-speed Internet services. Its primary function is to act as an endpoint device located at the user's premises. However, people are often confused as to whether an ONU can Work As a Router. In this blog post, we’ll explore What is ONU, What Is Router, ONU & Router’s main usage, and What is the difference between ONU and Router. Let’s take a closer look !
Part I. Meaning of...
| |
by: Hystou |
last post by:
Most computers default to English, but sometimes we require a different language, especially when relocating. Forgot to request a specific language before your computer shipped? No problem! You can effortlessly switch the default language on Windows 10 without reinstalling. I'll walk you through it.
First, let's disable language synchronization. With a Microsoft account, language settings sync across devices. To prevent any complications,...
|
by: Oralloy |
last post by:
Hello folks,
I am unable to find appropriate documentation on the type promotion of bit-fields when using the generalised comparison operator "<=>".
The problem is that using the GNU compilers, it seems that the internal comparison operator "<=>" tries to promote arguments from unsigned to signed.
This is as boiled down as I can make it.
Here is my compilation command:
g++-12 -std=c++20 -Wnarrowing bit_field.cpp
Here is the code in...
|
by: tracyyun |
last post by:
Dear forum friends,
With the development of smart home technology, a variety of wireless communication protocols have appeared on the market, such as Zigbee, Z-Wave, Wi-Fi, Bluetooth, etc. Each protocol has its own unique characteristics and advantages, but as a user who is planning to build a smart home system, I am a bit confused by the choice of these technologies. I'm particularly interested in Zigbee because I've heard it does some...
|
by: agi2029 |
last post by:
Let's talk about the concept of autonomous AI software engineers and no-code agents. These AIs are designed to manage the entire lifecycle of a software development project—planning, coding, testing, and deployment—without human intervention. Imagine an AI that can take a project description, break it down, write the code, debug it, and then launch it, all on its own....
Now, this would greatly impact the work of software developers. The idea...
|
by: isladogs |
last post by:
The next Access Europe User Group meeting will be on Wednesday 1 May 2024 starting at 18:00 UK time (6PM UTC+1) and finishing by 19:30 (7.30PM).
In this session, we are pleased to welcome a new presenter, Adolph Dupré who will be discussing some powerful techniques for using class modules.
He will explain when you may want to use classes instead of User Defined Types (UDT). For example, to manage the data in unbound forms.
Adolph will...
|
by: conductexam |
last post by:
I have .net C# application in which I am extracting data from word file and save it in database particularly. To store word all data as it is I am converting the whole word file firstly in HTML and then checking html paragraph one by one.
At the time of converting from word file to html my equations which are in the word document file was convert into image.
Globals.ThisAddIn.Application.ActiveDocument.Select();...
| |
by: TSSRALBI |
last post by:
Hello
I'm a network technician in training and I need your help.
I am currently learning how to create and manage the different types of VPNs and I have a question about LAN-to-LAN VPNs.
The last exercise I practiced was to create a LAN-to-LAN VPN between two Pfsense firewalls, by using IPSEC protocols.
I succeeded, with both firewalls in the same network. But I'm wondering if it's possible to do the same thing, with 2 Pfsense firewalls...
|
by: 6302768590 |
last post by:
Hai team
i want code for transfer the data from one system to another through IP address by using C# our system has to for every 5mins then we have to update the data what the data is updated we have to send another system
| |