☰
【C++】string的使用与模拟实现
2026/10/2 3:42:16 网站建设 项目流程

前言

std::string大概是 C++ 里被用得最多、也被误解得最多的类型。常见的误解有三个:第一,以为它就是char数组的语法糖,本质上和char buf[100]差不多;第二,以为std::string一定在堆上分配内存,所以"性能肯定比栈上的字符数组差";第三,以为"模拟实现"就是把std::string的源码抄一遍——实际上标准只规定了接口的行为,内部怎么存、有没有短字符串优化(Small String Optimization, SSO)、扩容因子是多少,全是实现定义的。

本文分两半:前半讲std::string的接口和真正会用到的用法,后半写一个"能编译、能跑、行为正确"的简化版MyString,通过它把"拷贝控制"这件事讲清楚。最后给出几个真会踩的坑,尤其是c_str()的悬垂指针和迭代器失效。

目标读者是刚学完 C 字符串、准备用std::string替换strcpy/strcat的人。本文代码以 C++17 为基准,GCC 13 / Clang 17 / MSVC 19.3x 均可编译。

一、std::string 到底是什么

std::string不是标准库里的"一个类",而是一个类型别名:

// 概念示意,不是可直接编译的声明 namespace std { template<class CharT, class Traits = char_traits<CharT>, class Allocator = allocator<CharT>> class basic_string; using string = basic_string<char>; }

也就是说,std::string是std::basic_string<char>。真正被标准规定的东西是basic_string的接口和复杂度要求;具体怎么实现由标准库厂商决定。你可以自己验证大小:

// C++17 #include <iostream> #include <string> #include <vector> int main() { std::cout << "sizeof(std::string) = " << sizeof(std::string) << '\n'; std::cout << "sizeof(std::vector<int>)= " << sizeof(std::vector<int>) << '\n'; std::string s = "hi"; std::cout << "size = " << s.size() << ", capacity = " << s.capacity() << '\n'; return 0; }

这段代码在 GCC 13(libstdc++,默认的 C++11 ABI)上通常打印sizeof(std::string) = 32。为什么是 32 而不是"一个指针加两个整数"的 24?因为多数实现在对象内部留了一小块本地缓冲区做 SSO:短字符串直接存在对象里,不碰堆。

SSO 的阈值是彻头彻尾的实现细节,标准一个字都没规定。下面是三家主流实现的常见情况,仅供理解,具体数值请以你本地的sizeof和capacity()实测为准:

实现所属编译器SSO 内部缓冲区常见容量sizeof(std::string)常见值
libstdc++(C++11 ABI,std::__cxx11::basic_string)GCC 5+15 个字符 + 结尾空字符32
libstdc++(旧 ABI,_GLIBCXX_USE_CXX11_ABI=0)GCC 5 之前无 SSO(写时复制)8
libc++Clang22 个字符 + 结尾空字符24
MSVC STLMSVC15 个字符 + 结尾空字符32

所以当有人问"std::string存 16 个字符会不会分配内存"时,正确答案是"取决于实现,GCC 的 libstdc++ 通常在 16 个字符时就已经转到堆上了,而 Clang 的 libc++ 要到 23 个字符才转"。别把它当标准断言。

二、真正会用到的接口

std::string的成员函数很多,但日常高频的其实就是下面这些。全部以标准的规定为准(需要精确签名时查 cppreference 或标准 [string] 一节)。

分类成员函数说明
容量size()/length()两者等价,返回字符个数(不含结尾空字符)
容量capacity()当前已分配空间能容纳多少字符
容量empty()是否为空
容量reserve(n)预留至少 n 个字符的空间,避免多次扩容
容量resize(n)改变元素个数,多出的位置用char()填充
容量clear()清空内容,容量一般不变
访问operator[](i)不检查越界,越界是 UB
访问at(i)越界抛std::out_of_range
访问front()/back()C++11 起;空串上调用是 UB
访问data()/c_str()返回指向内部缓冲的指针;c_str()一定以'\0'结尾
修改push_back(c)/pop_back()追加/删除末尾字符
修改append(s)/operator+=追加;+=比append更常用也更短
修改insert(pos, s)在 pos 处插入
修改erase(pos, n)删除从 pos 起 n 个字符
查找find(s, pos = 0)返回首次出现位置,找不到返回npos
查找rfind,find_first_of,find_first_not_of返回值同样用npos表示失败
截取substr(pos = 0, count = npos)返回新串(有拷贝),pos 越界抛std::out_of_range
比较compare(s)返回负值/0/正值

关于data()有一个必须记住的版本差异:

标准版本const std::string上的data()返回非 const 对象上的data()返回
C++11 / C++14const char*const char*
C++17 起const char*char*(可写)

也就是说,C++17 起非 const 的data()是可写的,但不能写超过size()的位置,也不能改写size()之后那个结尾空字符——那是 UB。

三、三个最常用的惯用法

用法一:拼接时先reserve。循环里反复+=会触发重新分配(reallocation):分配更大的块、把旧数据搬过去、释放旧块。搬一次就是 O(n)。如果提前知道大概长度,reserve能把这个开销压成一次。

// C++17 #include <string> #include <vector> #include <iostream> std::string join(const std::vector<std::string>& parts, char sep) { std::size_t total = 0; for (const auto& p : parts) total += p.size(); if (!parts.empty()) total += parts.size() - 1; std::string out; out.reserve(total); // 只预留一次 for (std::size_t i = 0; i < parts.size(); ++i) { if (i) out += sep; out += parts[i]; } return out; } int main() { std::vector<std::string> v{"alpha", "beta", "gamma"}; std::cout << join(v, ',') << '\n'; // alpha,beta,gamma return 0; }

用法二:用find+substr切分,注意npos的比较方式。std::string::npos是static const size_type npos = -1(即size_type的最大值)。比较时要小心类型宽度:把find的结果存进int会在 64 位平台上截断。

// C++17 #include <string> #include <vector> std::vector<std::string> split(const std::string& s, char sep) { std::vector<std::string> out; std::string::size_type start = 0; while (true) { std::string::size_type pos = s.find(sep, start); if (pos == std::string::npos) { out.push_back(s.substr(start)); break; } out.push_back(s.substr(start, pos - start)); start = pos + 1; } return out; }

用法三:需要 C 接口时用c_str(),但只在调用期间用。

// C++17 #include <string> #include <cstdio> int main() { std::string name = "cpp"; // 直接把指针交给 C 函数,printf 在本次调用内使用它,安全 std::printf("%s\n", name.c_str()); return 0; }

c_str()返回的指针在任何会修改这个 string 的操作之后都可能失效(包括push_back、+=、reserve导致的扩容)。这不是"可能失效"的模糊说法——标准规定这类操作会使指向元素的指针/引用失效的规则适用于data(),c_str()同理。

实战:一个可编译的简化版 MyString

下面这个类不追求接口完备,只实现最核心的部分:构造、析构、拷贝构造、拷贝赋值、移动构造、移动赋值、size、c_str、operator[]、operator+=,并且不实现 SSO(所有非空数据都在堆上)。通过它可以看清"拷贝控制"的五件事。

// C++17,单文件可直接编译:g++ -std=c++17 -Wall -Wextra mystring.cpp #include <cstddef> #include <cstring> #include <iostream> #include <utility> class MyString { public: // 默认构造:空串也要有一个合法的 buf_,保证 c_str() 可用 MyString() : size_(0), cap_(0), buf_(new char[1]) { buf_[0] = '\0'; } MyString(const char* s) { size_ = std::strlen(s); cap_ = size_; buf_ = new char[size_ + 1]; std::memcpy(buf_, s, size_ + 1); // 连结尾 '\0' 一起拷 } // 拷贝构造:深拷贝 MyString(const MyString& other) { size_ = other.size_; cap_ = other.size_; buf_ = new char[size_ + 1]; std::memcpy(buf_, other.buf_, size_ + 1); } // 移动构造:接管别人的缓冲区,并把对方置为有效但空的状态 MyString(MyString&& other) noexcept : size_(other.size_), cap_(other.cap_), buf_(other.buf_) { other.size_ = 0; other.cap_ = 0; other.buf_ = new char[1]; // 对方仍然要能析构、能 c_str() other.buf_[0] = '\0'; } MyString& operator=(const MyString& other) { if (this != &other) { // 必须自赋值检查,否则自己释放自己 MyString tmp(other); // 先拷贝,异常安全 swap(tmp); } return *this; } MyString& operator=(MyString&& other) noexcept { if (this != &other) { MyString tmp(std::move(other)); swap(tmp); } return *this; } ~MyString() { delete[] buf_; } void swap(MyString& other) noexcept { std::swap(size_, other.size_); std::swap(cap_, other.cap_); std::swap(buf_, other.buf_); } std::size_t size() const { return size_; } bool empty() const { return size_ == 0; } const char* c_str() const { return buf_; } char& operator[](std::size_t i) { return buf_[i]; } // 不检查越界 const char& operator[](std::size_t i) const { return buf_[i]; } void reserve(std::size_t n) { if (n <= cap_) return; char* nb = new char[n + 1]; std::memcpy(nb, buf_, size_ + 1); delete[] buf_; buf_ = nb; cap_ = n; } MyString& operator+=(const MyString& rhs) { if (rhs.size_ == 0) return *this; if (size_ + rhs.size_ > cap_) { reserve(size_ + rhs.size_); // 简化策略:需要多少要多少 } std::memcpy(buf_ + size_, rhs.buf_, rhs.size_ + 1); size_ += rhs.size_; return *this; } MyString& operator+=(const char* rhs) { return *this += MyString(rhs); } private: std::size_t size_; std::size_t cap_; char* buf_; }; MyString operator+(MyString lhs, const MyString& rhs) { lhs += rhs; // 传值 + 返回,天然享受移动语义 return lhs; } std::ostream& operator<<(std::ostream& os, const MyString& s) { return os << s.c_str(); } int main() { MyString a("Hello"); MyString b = a; // 拷贝构造 b += MyString(", world"); std::cout << b << " (size=" << b.size() << ")\n"; MyString c = a + b; // 移动构造 std::cout << c << " (size=" << c.size() << ")\n"; MyString d; d = std::move(c); // 移动赋值 std::cout << d << '\n'; a = a; // 自赋值,必须是安全的 std::cout << a << '\n'; return 0; }

这段代码在 GCC 13 / Clang 17 上用-std=c++17 -Wall -Wextra编译应无警告。几个设计点值得说明:


  1. 默认构造也分配 1 字节,这样c_str()永远返回一个合法的、以'\0'结尾的指针。

  2. 移动构造里noexcept不是装饰。标准容器(比如std::vector)在扩容时判断元素的移动构造函数是否noexcept:是则用移动,否则为了强异常保证会退回到拷贝。标上noexcept能让容器选择更省的路径。

  3. 移动后源对象仍处于"有效但未指定"状态。我在移动构造里给源对象重新分配了空缓冲区,这比"留一个空指针"更安全(后者会让源对象的c_str()直接崩溃)。

  4. 赋值用"拷贝并交换"(copy-and-swap),天然处理自赋值,且异常安全。

  5. operator+按值接收左操作数,于是lhs本身就是一份拷贝,可以直接在上面追加再返回——返回时享受移动语义。


需要坦白一点:这个简化版的移动构造里做了一次new char[1],却把移动构造标成了noexcept。一旦这次分配失败抛出std::bad_alloc,程序会直接调用std::terminate。真实的标准库实现靠 SSO 或共享的空串静态对象避免这次分配,它们的移动构造是真的不会失败。写生产代码时不要照抄这一点:要么让移动构造真的不做可能抛异常的事,要么就别标noexcept——而不标noexcept又会失去容器扩容时优先移动的机会,这正是 SSO 重要性的来源之一。

常见坑点

坑 1:把c_str()的返回值存下来跨语句使用。

❌

const char* p = s.c_str(); s += "x"; // 可能触发扩容,p 变成悬垂指针 std::printf("%s\n", p); // UB,标准不保证任何行为

✅ 只在同一个表达式/同一个调用内用c_str();确实要留住,就复制一份到std::string或std::vector<char>。

坑 2:substr越界不会报错,但at会。

❌std::string s = "abc"; char c = s[10];——operator[]不检查,越界是 UB。

✅ 用s.at(10),它在越界时抛std::out_of_range。注意substr(pos)在pos > size()时也抛std::out_of_range。

坑 3:把npos塞进int。

❌

int pos = s.find("x"); // npos 在 64 位平台被截断,判断结果错乱 if (pos == std::string::npos) // 类型不匹配,比较结果不可靠

✅ 用std::string::size_type(或auto)接收find的返回值,再和std::string::npos比较。

坑 4:以为reserve之后指针就永远稳定了。

❌ 在循环里reserve一次,然后一路追加并缓存data()返回的指针。

✅reserve只是减少扩容次数;任何可能改动size或capacity的操作都可能让之前的指针失效。要长期持有就存下标,不存指针。

坑 5:中文字符串用size()当"字数"。

❌std::string s = "中文";然后认为s.size() == 2。UTF-8 下一个汉字通常是 3 字节,因此s.size()是 6。

✅ 明确区分"字节数"和"字符数"。字符级处理需要宽字符、std::u8string(C++20)或第三方 Unicode 库;只想知道 UTF-8 码点个数,可以手写一个按首字节判断续字节个数的计数器。

坑 6:循环里s = s + "x";而不是s += "x";。

❌

for (int i = 0; i < 1000; ++i) s = s + "x"; // 每次都构造临时串再拷贝赋值

✅s += "x";——operator+=直接在原串上追加,不需要构造完整副本。

总结

主题结论
std::string的本质std::basic_string<char>的别名,接口由标准规定,布局由实现决定
SSO实现细节,GCC 的 libstdc++ 与 MSVC STL 常见 15 字符,Clang 的 libc++ 常见 22 字符
取字符指针用c_str(),且只在调用期间使用;C++17 起非 const 的data()可写
查找失败返回std::string::npos,必须用std::string::size_type接收
效率循环拼接前先reserve;用+=而不是s = s + ...
实现自己的字符串类必须成套实现拷贝构造/拷贝赋值/移动构造/移动赋值/析构,移动构造标noexcept

std::string的接口看着平易近人,真正的难点在两个地方:一是生命周期——所有返回指针的函数(c_str、data)都把"什么时候会失效"的责任交给了调用者;二是实现差异——SSO、容量增长策略这些看起来"应该是标准"的东西其实全是厂商自由发挥。把这两点记牢,用std::string就很少会出问题;自己写一个简化版,则是把拷贝控制这套规则真正内化的最快办法。

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询