[vb5] Ascii Character Set (128 – 255)

Pagina: 1
Acties:

  • Inge801
  • Registratie: Januari 2002
  • Niet online

Inge801

Iron Maiden

Topicstarter
http://picserver.student.utwente.nl/getpicture.php?id=70710

Ik ben in VB5 een programma aan het maken, maar vb5 hanteert zoals in mijn link te zien is een hele rare Ascii table? Zo staan er af en toe in mijn strings hele rare karakters.

code:
1
2
debug.Print Asc("ë")
 235


terwijl ik hier toch graag 137 uit had. Iemand een idee?

你还记得吗 记忆的炎夏


  • drm
  • Registratie: Februari 2001
  • Laatst online: 09-06-2025

drm

f0pc0dert

235 is de ë in de unicode (UTF-8 ?) tabel, niet in ascii... Misschien moet je eens kijken welke encoding er gehanteerd wordt :?

Music is the pleasure the human mind experiences from counting without being aware that it is counting
~ Gottfried Leibniz


  • Inge801
  • Registratie: Januari 2002
  • Niet online

Inge801

Iron Maiden

Topicstarter
Hmm ok, stel dat het Unicode is, en ik heb een ascii string, weet iemand hoe ik mijn string dan converteer zodat vb hem wel correct kan weergeven?

你还记得吗 记忆的炎夏


  • Limhes
  • Registratie: Oktober 2001
  • Laatst online: 19-08 19:06
Controls for Unicode development in Visual Basic

[ Voor 6% gewijzigd door Limhes op 15-03-2003 15:35 ]


  • Inge801
  • Registratie: Januari 2002
  • Niet online

Inge801

Iron Maiden

Topicstarter
Ik denk dat het probleem bij Winsock ligt, ik heb 6.0 sp5, daar lijkt het verzenden van unicode strings mis te gaan, zodra ik er een verzend komt er een verbindingserror.

Heet iemand misschien tips hiervoor?

edit: http://www.qbssoftware.com/ hiermee lijkt het te kunnen, alleen dat kan ik niet betalen :|

[ Voor 20% gewijzigd door Inge801 op 18-03-2003 14:04 ]

你还记得吗 记忆的炎夏


  • MSalters
  • Registratie: Juni 2001
  • Laatst online: 21-08 17:14
ASCII is 7 bits, 0-127. Engels heeft geen diakrieten. De vraag is dus wat je dan wel hebt, google eens naar ISO-8859. Dat zijn 8-bits supersets van ASCII.

Man hopes. Genius creates. Ralph Waldo Emerson
Never worry about theory as long as the machinery does what it's supposed to do. R. A. Heinlein


  • Limhes
  • Registratie: Oktober 2001
  • Laatst online: 19-08 19:06
En unicode is 2 bytes per karakter (of vergis ik me nu?), dus zou je combinaties van twee op elkaar volgende bytes kunnen omzetten naar zijn ascii equivalent.

  • Limhes
  • Registratie: Oktober 2001
  • Laatst online: 19-08 19:06
Je kan de string gewoon omzetten naar een aaneenschakeling van bytes, die je vervolgens wel kan overzenden met winsock. Hierna zet je ze weer om naar een string:


Visual Basic:
1
2
3
4
5
6
7
8
Dim bytes() As Byte
Dim unicode As String

' string naar bytes:
bytes = unicode & vbNullChar

' bytes naar string:
unicode = CStr(bytes)

[ Voor 3% gewijzigd door Limhes op 18-03-2003 15:41 ]


  • MSalters
  • Registratie: Juni 2001
  • Laatst online: 21-08 17:14
Unicode is 20 bits, sinds Unicode 3.1. Maar zelfs voor 16 bits heb je dus 3 ASCII karakters = 21 bits nodig. 2 ASCII karakters = 14 bits.
(En dan heb ik het nog niet over de gewoonte van sommige ASCII libraries om ASCII NUL als het einde van een string te zien)

Om het ingewikkelder te maken is de MSWin encoding van Unicode een 16/32 bits encoding, sommige karakters zijn dus 2 bytes en andere 4.

Alsof dat nog niet ingewikkeld genoeg is bevatten die karakters dus regelmatig NUL-bytes: karakter 0x0400 is dus de bytes 0x04 en 0x00. TCP/IP en WinSock zijn byte-oriented, dus als die per ongeluk ASCII tekst denken te zien dan gaat het dus fout.

De oplossing voor dat probleem is om de Unicode als UTF-8 over WinSock te zenden.

Man hopes. Genius creates. Ralph Waldo Emerson
Never worry about theory as long as the machinery does what it's supposed to do. R. A. Heinlein


  • Limhes
  • Registratie: Oktober 2001
  • Laatst online: 19-08 19:06
MSalters schreef op 18 maart 2003 @ 15:42:
...
Om het ingewikkelder te maken is de MSWin encoding van Unicode een 16/32 bits encoding, sommige karakters zijn dus 2 bytes en andere 4.
...
Visual Basic maakt idd 2 bytes van de meeste karakters, maar hoe weet je nu of er twee combinaties van 2 bij elkaar horen? Dan kunnen die dus per 2 bytes zelf niets betekenen want anders is het onomkeerbaar.
Leg es uit...

  • .oisyn
  • Registratie: September 2000
  • Laatst online: 22-08 13:19

.oisyn

Moderator Devschuur®

Demotivational Speaker

bedoel je hoe het systeem weet dat de volgende char er een van 4 bytes is ipv 2? Ik ben niet zo bekend met character encodings, maar waarschijnlijk bestaat een 2-byte unicode char gewoon uit 15 bits bits. Als de 16e bit van die char 1 is wil dat zeggen dat er nog 2 komen, en dan heb je uiteindelijk een char van 31 bits

Ik zeg dus niet dat het zo gaat, maar gewoon dat het waarschijnlijk ongeveer zo gaat ;)

Give a man a game and he'll have fun for a day. Teach a man to make games and he'll never have fun again.


  • Limhes
  • Registratie: Oktober 2001
  • Laatst online: 19-08 19:06
.oisyn schreef op 18 March 2003 @ 15:55:
...waarschijnlijk ongeveer...
Lijkt me dus beter dat je nog even nazoekt Vincent (Inge)

Nofi Oisyn :P (krijgt Orbb trouwens built-in Unicode functionaliteit?)

  • .oisyn
  • Registratie: September 2000
  • Laatst online: 22-08 13:19

.oisyn

Moderator Devschuur®

Demotivational Speaker

denk het niet, maar wie weet ;)

Give a man a game and he'll have fun for a day. Teach a man to make games and he'll never have fun again.


  • MSalters
  • Registratie: Juni 2001
  • Laatst online: 21-08 17:14
Unicode site over UTF-16 ( UTF-16 is wat MS gebruikt )
Q: What is UTF-16?

A: Unicode was originally designed as a pure 16-bit encoding, aimed at representing all modern scripts. (Ancient scripts were to be represented with private-use characters.) Over time, and especially after the addition of over 14,500 composite characters for compatibility with legacy sets, it became clear that 16-bits were not sufficient for the user community. Out of this arose UTF-16.

UTF-16 allows access to 63K characters as single Unicode 16-bit units. It can access an additional 1M characters by a mechanism known as surrogate pairs. Two ranges of Unicode code values are reserved for the high (first) and low (second) values of these pairs. Highs are from 0xD800 to 0xDBFF, and lows from 0xDC00 to 0xDFFF. In Unicode 3.0, there are no assigned surrogate pairs. Since the most common characters have already been encoded in the first 64K values, the characters requiring surrogate pairs will be relatively rare (see below). [MD]

Q: Does UTF-16 have an alternative representation?

A: Yes, all characters represented in UTF-16, both those represented with 16 bits and those with a surrogate pair, can be represented as a single 32-bit unit in UTF-32. This single 4 code unit corresponds to the Unicode scalar value, which is the abstract number associated with a Unicode character. UTF-32 is a subset of the encoding mechanism called UCS-4 in ISO 10646

Man hopes. Genius creates. Ralph Waldo Emerson
Never worry about theory as long as the machinery does what it's supposed to do. R. A. Heinlein

Pagina: 1